Does a batch API cut the cost of grading agent traces with an LLM judge?
Yes. Anthropic’s Message Batches API prices standard chat completions at half their normal per-token rate, in exchange for an asynchronous response instead of an instant one; Anthropic’s docs put typical turnaround under an hour, with a stated ceiling of 24. That trade suits LLM judge grading well, since a trace almost never needs to be graded the instant it lands. A nightly grading run already accepts a delay before its verdict is useful, so moving it onto a batch endpoint costs nothing beyond a wait that was already built into the schedule. Prompt caching and batching stack on top of each other, too: grading many traces that share context in one batch job earns the cache discount and the batch discount together, so a pipeline paying for both trims more off the bill than either discount alone.