all answers

Agent reliability

Graders

A grader is a check that reads a trace of an agent's behavior and returns a verdict about one specific thing the agent was supposed to do. Running one grader across many traces turns a general sense of quality into a rate for one behavior, and a rate can be compared across agent versions and over time.

Graders come in three forms, ordered by cost per verdict. Deterministic checks are code that inspects the trace, cheap enough to run on every output. Trained classifiers read text and score a single property at low cost. LLM judges use a model to reason about the output. They cost the most per verdict and cover judgments that require reasoning about meaning.

Two properties determine what a grader's verdict means. The first is scope. A grader tied to one observable behavior, with a stated condition for when it applies, gives a verdict that says exactly which behavior failed. As scope widens toward overall quality, each verdict carries less information about what actually happened. The second is origin. A grader written from a failure observed in real traffic measures something that actually occurs. A grader written from speculation measures what its author imagined, which may differ from what the agent does in practice.

Graders also get stale when the agent's intent changes. Each grader encodes what the agent was supposed to do when it was written, so when the intent changes, the grader keeps judging against the old one. It still returns verdicts, so its rate can move for reasons unrelated to what the agent is now meant to do.

12 questions

Answered, plainly.

Does a regression grader have to be an LLM judge?No. A deterministic check that asserts the exact condition that broke is cheaper and more stable than an LLM judge, which only earns its cost on failures that need judging meaning.answer →What stops a description from producing a grader that measures the wrong thing?Nothing by itself. A description only encodes what its author imagined the failure looks like, and needs grounding in a real trace where that failure actually happened before it ships.answer →How do you know when a grader no longer matches what the agent does?Not from the pass rate. Check whether the prompt, tool contract, or output schema the grader was written against has changed since, because it keeps judging the old version either way.answer →What does an eval score actually claim about my agent?It's a claim about the eval set the graders ran on, not the agent as a whole; it says nothing about a failure mode no grader in the set checks for.answer →What is a grader?A grader is a check that reads a trace of an agent's behavior and returns a verdict about one specific thing the agent was supposed to do.answer →How do you measure whether a grader itself is accurate?Run it against a human-labeled set with a known right answer and compare its verdicts to those labels; the resulting precision and recall are its accuracy.answer →How do you catch a security flaw in AI-generated code that still runs correctly?Run the code against its test suite and separately scan the diff for known vulnerability patterns; code that works can still be insecure.answer →Why do vision agents need a grader that checks perception and action separately?Because an agent can describe an image correctly and act on the wrong part of it, or the reverse, and a single verdict can't tell you which.answer →How do you score turn-taking smoothness when a user interrupts the agent?With a dedicated grader on the turn itself; a transcript-only eval can score every reply as accurate while the caller was talked over or lost.answer →Why isn't grading an agent's final answer enough to catch a multi-step failure?Because a wrong intermediate step can still produce a final reply that reads fine, so grading only the end catches the symptom, not the step that caused it.answer →Can a security-pattern grader replace a SAST scan?No. A pattern grader reads the agent's diff for known vulnerabilities; a SAST scan reads the whole repo. They check different surfaces, and neither replaces the other.answer →Is an agent that correctly refuses a task a completion failure?To a task-completion grader, yes: a required refusal and an abandoned task both show no state change, and the grader can't tell them apart.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y