# Does a batch API cut the cost of grading agent traces with an LLM judge?

Yes. Anthropic's Message Batches API prices standard chat completions at half their normal per-token rate, in exchange for an asynchronous response instead of an instant one; Anthropic's docs put typical turnaround under an hour, with a stated ceiling of 24. That trade suits LLM judge grading well, since a trace almost never needs to be graded the instant it lands. A [nightly grading run](/answers/eval-costs/which-graders-belong-on-every-pr-and-which-on-nightly) already accepts a delay before its verdict is useful, so moving it onto a batch endpoint costs nothing beyond a wait that was already built into the schedule. Prompt caching and batching stack on top of each other, too: grading many traces that share context in one batch job earns the cache discount and the batch discount together, so a pipeline paying for both trims more off the bill than either discount alone.

---

Sources:
- Anthropic, Batch processing documentation: https://platform.claude.com/docs/en/build-with-claude/batch-processing (fetched 2026-09-03)

Source: https://tessary.ai/answers/eval-costs/does-a-batch-api-cut-llm-judge-grading-costs
More on Eval costs: https://tessary.ai/answers/eval-costs
From Tessary, agent reliability for AI agents in production: https://tessary.ai
