answers
One question, one page.
48 concepts, 394 questions answered. Each page answers one question directly in its first paragraph, then shows the work and cites where the numbers came from.
11 concepts
Agent reliability
Agent reliabilityAgent reliability is the degree to which an agent's behavior in production stays consistent with what it was built to do, across the inputs it actually receives and over time as everything around it changes.10 answers →Cause attributionCause attribution is the step from "quality dropped" to "this specific change caused it." Detection establishes that a regression happened; attribution names the change responsible.8 answers →Deploy gatesA deploy gate is a check that runs a set of evaluations against a change before it merges and blocks the merge on failure.7 answers →Eval costsEval costs are the arithmetic of judging an agent's traffic.7 answers →Eval datasetsAn eval dataset is the set of cases an agent gets judged against: inputs paired with the behavior expected on them.10 answers →Failure replayFailure replay is the practice of reproducing a production failure with its full original context, so that a proposed fix can be verified against the case that actually happened.8 answers →GradersA grader is a check that reads a trace of an agent's behavior and returns a verdict about one specific thing the agent was supposed to do.12 answers →InstrumentationInstrumentation is the code that records what an agent does while it runs and emits that record as telemetry.10 answers →LLM as judgeLLM-as-judge is a prompt that grades an agent's output by reasoning over it against one or more specific intents, the things the agent was supposed to do.17 answers →Regression detectionA regression is a drop in agent quality caused by a change.9 answers →Silent failuresA silent failure is an agent run that completes normally and produces a wrong result.6 answers →
6 concepts
General agent concepts
Agent harnessAn agent harness is everything in an agent that isn't the model: the system prompt and instruction files, the tools and their descriptions, how context is assembled and trimmed, hooks and middleware, any sandbox or subagents, and the loop that calls the model, runs what it asks for, and decides when to stop.8 answers →Failure modesA failure mode is a recurring, nameable way an agent goes wrong.14 answers →HallucinationsA hallucination is an agent stating something its sources don't support.11 answers →Retrieval augmented generationRetrieval-augmented generation, RAG, is the pattern of fetching relevant content at request time and putting it in the model's context before it answers.11 answers →System one modelsA System One model is a model built to make a decision rather than write text.9 answers →Tool callingTool calling is how a language model acts outside its own text.15 answers →
14 concepts
Tessary
Tessary behavior driftbehavior_drift is one of Tessary's built-in classifiers.7 answers →Tessary call sitesA call site is a place in your code where an agent with a specific function runs.4 answers →Tessary casesA case is Tessary's unit of investigation.8 answers →Tessary classifiersTessary's classifiers are the cheap checks that run on every production trace.8 answers →Tessary cost driftcost_drift is one of Tessary's built-in classifiers: a tripwire that watches each call site's spend against its own past.5 answers →Tessary custom classifiersCustom classifiers are classifiers Tessary fine-tunes for your agent.1 answer →Tessary duration driftduration_drift is one of Tessary's built-in classifiers.5 answers →Tessary frustrationFrustration is one of Tessary's built-in classifiers.10 answers →Tessary groundednessGroundedness is one of Tessary's built-in classifiers.7 answers →Tessary malformed outputmalformed_output is one of Tessary's built-in classifiers, a deterministic one: checks that pass or fail, no judgment involved.6 answers →Tessary RCARCA is the explain step at the end of Tessary's pipeline.8 answers →Tessary secret leaksecret_leak is one of Tessary's built-in classifiers, a deterministic credential detector.4 answers →Tessary tool errortool_error is one of Tessary's built-in classifiers.5 answers →Tessary tracesA trace is one unit in Tessary: the stored record of one agent turn.9 answers →
11 concepts
Research and benchmarks
Agent benchmarksAn agent benchmark is a fixed set of tasks, an environment to run them in, and a scoring rule, used to compare models and agent designs on something repeatable.8 answers →Agent loops research"When Agents Do Not Stop" (arXiv:2607.01641) is a study of agents that never finish.5 answers →BfclBFCL, the Berkeley Function Calling Leaderboard, tests whether a model can use tools correctly.5 answers →CursorbenchCursorBench is Cursor's own test of coding agents, built from real tasks that Cursor users gave the agent in their editor.5 answers →GaiaGaia2 is Meta's test for personal-assistant agents, released in September 2025.4 answers →Judge bias researchAn LLM judge is a model that scores another model's output.8 answers →Mast taxonomyMAST is a taxonomy of how multi-agent LLM systems fail, from the Berkeley paper "Why Do Multi-Agent LLM Systems Fail?" (arXiv:2503.13657).7 answers →ReliabilitybenchReliabilityBench (arXiv:2601.06112) is a benchmark for tool-using agents that measures reliability instead of a single-run success rate.6 answers →Swe benchSWE-bench Pro is a coding test run by Scale AI, released in September 2025.7 answers →Tau benchtau-bench is Sierra's test for customer-service agents.5 answers →Terminal benchTerminal-Bench gives an AI a computer terminal and a job to finish.5 answers →
6 concepts
Frameworks and tooling
Claude agent SDK evalsThe Claude Agent SDK is the harness behind Claude Code, published so you can build your own agent on the same loop: gather context, act, check the result, repeat.17 answers →JevJev is TypeSafe AI's decision model, launched September 15, 2026, and the first sold as a System One model.9 answers →Langgraph evalsLangGraph is LangChain's library for building an agent as a graph.9 answers →Mcp tool reliabilityThe Model Context Protocol (MCP) is a standard way to give a model tools that live in a separate process or service.10 answers →Openai agents SDK evalsThe OpenAI Agents SDK is a small runtime for building agents.16 answers →Otel genai conventionsThe OpenTelemetry GenAI semantic conventions are a shared vocabulary for describing an LLM call, a tool call, or an agent step as a span.9 answers →
Two ways to run Tessary.
Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.
Tessary Cloud
We host it for you. Send your first trace with nothing to deploy and no model key.
what's included
- traces
- 10,000 per calendar month
- stored trace data
- 1 GB
- retention
- 30 days
- model credit
- $10, one-time, for triage and root-cause analysis
- credit card
- not required
Self-hosted Tessary
Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.
Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md
docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y