all answers

Research and benchmarks

Judge bias research

An LLM judge is a model that scores another model's output. Judges have biases that push every verdict the same way, and 2026 research has measured them across enough models to say what matters.

The biggest one is style. "Judging the Judges" (arXiv:2604.23178) tested five judges from four providers and found they favour markdown-formatted answers over plain prose by a wide margin, far more than they favour whichever answer comes first. Length bias depends on the judge: some prefer longer answers, Claude prefers shorter ones, and GPT-4o was neutral. The same study found a mid-tier model with the right debiasing prompt matched or beat frontier judges at about a fifteenth of the cost.

Self-preference, where a judge favours its own model's output, does not go away with capability. A study across 20 models (arXiv:2604.22891) found stronger models were no less biased, and sometimes more. A structured, multi-dimension scoring prompt cut it by about a third.

Two findings change how you should validate a judge. The largest evaluation to date, 21 judges and about 541,000 judgments (arXiv:2606.19544), showed that raw agreement with humans overstates judge quality by 33 to 41 points once you correct for chance, and that judge rankings move by up to 14 places from one benchmark to the next. And on agent tool-calling traces specifically, AgentJudgeBench (arXiv:2608.26623) found every judge tested, small or frontier, landed in the same 77 to 82% agreement band on hard tasks without ground truth. Chain-of-thought and temperature made no difference; a structured rubric helped by up to 6.5 points.

So: measure your judge against your own labelled traces, correct for chance, and expect a ceiling on hard cases that a bigger model won't lift.

8 questions

Answered, plainly.

How accurate are LLM judges on agent tool-calling traces?On hard tool-calling traces with no ground truth most judges plateau at 77 to 82% agreement regardless of scale, though weak or strong agents shift that band.answer →Do LLM judges score markdown-formatted answers higher than plain prose?Yes, and it's the largest bias measured: style bias runs 0.10 to 0.76 across judges, far past the 0.04 ceiling for favoring whichever answer came first.answer →Does a bigger model make a better judge?No. A 20-model study found self-preference bias uncorrelated with capability, and a separate 541,000-judgment study found scale doesn't stabilize rankings.answer →Does a judge that ranks well on one benchmark rank well on another?Not reliably: the largest LLM-judge evaluation to date found rankings shift by up to 14 positions when the same judges are scored on a different benchmark.answer →Does a more capable LLM judge show less self-preference bias?No. A 20-model study found more capable LLM judges show no less self-preference bias, sometimes more; a restructured scoring prompt cut it 31.5%.answer →Do different LLM judges agree when scoring the same rubric?Barely. A 2026 single-author study found near-zero agreement between judges scoring SEO content rubrics, Krippendorff's alpha of 0.042, though each judge scored consistently with itself.answer →How do I tell whether my judge is biased?Swap the order of what it's comparing and check how far its rate of picking 'A' strays from 50%; run-to-run consistency alone won't catch a judge that's just consistently biased.answer →Does a judge panel's verdict depend on which model families sit on it?Yes. A controlled study found swapping which model families sit on a judge panel changed 18.5% of its verdicts, since judges favor their own family's answers.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y