answers

One question, one page.

48 concepts, 394 questions answered. Each page answers one question directly in its first paragraph, then shows the work and cites where the numbers came from.

11 concepts

Agent reliability

Agent reliabilityAgent reliability is the degree to which an agent's behavior in production stays consistent with what it was built to do, across the inputs it actually receives and over time as everything around it changes.10 answers →Cause attributionCause attribution is the step from "quality dropped" to "this specific change caused it." Detection establishes that a regression happened; attribution names the change responsible.8 answers →Deploy gatesA deploy gate is a check that runs a set of evaluations against a change before it merges and blocks the merge on failure.7 answers →Eval costsEval costs are the arithmetic of judging an agent's traffic.7 answers →Eval datasetsAn eval dataset is the set of cases an agent gets judged against: inputs paired with the behavior expected on them.10 answers →Failure replayFailure replay is the practice of reproducing a production failure with its full original context, so that a proposed fix can be verified against the case that actually happened.8 answers →GradersA grader is a check that reads a trace of an agent's behavior and returns a verdict about one specific thing the agent was supposed to do.12 answers →InstrumentationInstrumentation is the code that records what an agent does while it runs and emits that record as telemetry.10 answers →LLM as judgeLLM-as-judge is a prompt that grades an agent's output by reasoning over it against one or more specific intents, the things the agent was supposed to do.17 answers →Regression detectionA regression is a drop in agent quality caused by a change.9 answers →Silent failuresA silent failure is an agent run that completes normally and produces a wrong result.6 answers →

6 concepts

General agent concepts

14 concepts

Tessary

11 concepts

Research and benchmarks

6 concepts

Frameworks and tooling

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y