How accurate are LLM judges on agent tool-calling traces?
On hard tool-calling traces with no ground truth, most tested judges plateau in the same 77 to 82% agreement band no matter their scale, per AgentJudgeBench’s six judges, ranging from 20 billion parameters to frontier scale, grading 3,808 tool-calling records. That ceiling isn’t universal, though: traces from the weakest agent model tested scored noticeably lower, in the high 60s to low 70s, and traces from the strongest open model scored higher, in the mid 80s, so the plateau describes a typical case rather than every one. Judges also do better on easier queries across the board, by roughly five to seven points, since difficulty here means how ambiguous the request is, not how hard the tool call itself is.
A structured rubric prompt helped one judge-generator pairing by up to 6.5 points and barely moved another, so it’s worth testing against your own labeled traces rather than assuming it always works. Chain-of-thought reasoning and temperature changed almost nothing either way.