Does a batch API cut the cost of grading agent traces with an LLM judge?

Yes. Anthropic’s Message Batches API prices standard chat completions at half their normal per-token rate, in exchange for an asynchronous response instead of an instant one; Anthropic’s docs put typical turnaround under an hour, with a stated ceiling of 24. That trade suits LLM judge grading well, since a trace almost never needs to be graded the instant it lands. A nightly grading run already accepts a delay before its verdict is useful, so moving it onto a batch endpoint costs nothing beyond a wait that was already built into the schedule. Prompt caching and batching stack on top of each other, too: grading many traces that share context in one batch job earns the cache discount and the batch discount together, so a pipeline paying for both trims more off the bill than either discount alone.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y