all answers

Agent reliability

Eval costs

Eval costs are the arithmetic of judging an agent's traffic. Grading methods differ in per-event cost by orders of magnitude, and at scale those differences dominate the bill. A deterministic check costs effectively nothing per event. Grading with a language model pays for inference on every event, and that price varies widely with the model chosen.

The gap comes from where costs land. Some methods carry their cost per event: every trace a model judges incurs an inference charge that recurs with traffic. Others carry their cost up front: a trained classifier's cost is paid once at training time, so grading an additional trace with it costs close to nothing. Because the per-event methods carry nearly all the marginal cost, total spend tracks the share of traffic that reaches them multiplied by traffic volume.

This arithmetic decides what evaluation is feasible. Low-probability failures surface only when a large share of traffic gets examined, so the per-event cost of grading caps how much traffic can be examined, and with it which failures can be observed at all.

7 questions

Answered, plainly.

Does reading every trace cost the same as grading every trace?No. Reading a trace is deterministic, near-free CPU work; grading it with a language model is a paid inference call that scales with volume.answer →Does running several graders on one trace cost more than running one?Not proportionally, if the graders share a prompt cache: input cost runs about 1.15 + 0.1N times the base rate for N graders, not N times it.answer →What drives the cost of evaluating an AI agent?Cost is driven by which grading methods your traffic hits: deterministic checks cost almost nothing per event, a language model judging a trace does, every time.answer →Which graders belong on every PR and which belong on a nightly run?Cheap, deterministic graders belong on every PR; anything that calls a language model belongs on a scheduled run, because that call is priced per commit.answer →Why do quality layers built on an LLM judge end up sampling instead of grading everything?A generative judge makes a paid call on every trace it grades, so the bill scales with traffic, and sampling is the cheapest lever to bring it back down.answer →Does a batch API cut the cost of grading agent traces with an LLM judge?Yes. Anthropic's Message Batches API cuts per-token cost in half for grading calls that can wait, the same tradeoff a nightly judge run already accepts.answer →Why does a classifier cost less to run than an LLM judge?A narrow classifier does less work per trace than a generative judge. Training happens once, but checking each new trace still costs compute.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y