# Judge bias research

An LLM judge is a model that scores another model's output. Judges have biases that push every
verdict the same way, and 2026 research has measured them across enough models to say what
matters.

The biggest one is style. "Judging the Judges" (arXiv:2604.23178) tested five judges from four
providers and found they favour markdown-formatted answers over plain prose by a wide margin, far
more than they favour whichever answer comes first. Length bias depends on the judge: some prefer
longer answers, Claude prefers shorter ones, and GPT-4o was neutral. The same study found a
mid-tier model with the right debiasing prompt matched or beat frontier judges at about a
fifteenth of the cost.

Self-preference, where a judge favours its own model's output, does not go away with capability.
A study across 20 models (arXiv:2604.22891) found stronger models were no less biased, and
sometimes more. A structured, multi-dimension scoring prompt cut it by about a third.

Two findings change how you should validate a judge. The largest evaluation to date, 21 judges
and about 541,000 judgments (arXiv:2606.19544), showed that raw agreement with humans overstates
judge quality by 33 to 41 points once you correct for chance, and that judge rankings move by up
to 14 places from one benchmark to the next. And on agent tool-calling traces specifically,
AgentJudgeBench (arXiv:2608.26623) found every judge tested, small or frontier, landed in the
same 77 to 82% agreement band on hard tasks without ground truth. Chain-of-thought and
temperature made no difference; a structured rubric helped by up to 6.5 points.

So: measure your judge against your own labelled traces, correct for chance, and expect a
ceiling on hard cases that a bigger model won't lift.

## Questions answered under this concept

- [Do different LLM judges agree when scoring the same rubric?](https://tessary.ai/answers/judge-bias-research/do-different-llm-judges-agree-when-scoring-the-same-rubric)
- [Does a judge that ranks well on one benchmark rank well on another?](https://tessary.ai/answers/judge-bias-research/do-llm-judge-rankings-transfer-across-benchmarks)
- [Do LLM judges score markdown-formatted answers higher than plain prose?](https://tessary.ai/answers/judge-bias-research/do-llm-judges-score-markdown-formatted-answers-higher)
- [Does a bigger model make a better judge?](https://tessary.ai/answers/judge-bias-research/does-a-bigger-model-make-a-better-judge)
- [Does a more capable LLM judge show less self-preference bias?](https://tessary.ai/answers/judge-bias-research/does-a-more-capable-llm-judge-show-less-self-preference-bias)
- [Does a judge panel's verdict depend on which model families sit on it?](https://tessary.ai/answers/judge-bias-research/does-judge-panel-composition-change-the-verdict)
- [How accurate are LLM judges on agent tool-calling traces?](https://tessary.ai/answers/judge-bias-research/how-accurate-are-llm-judges-on-tool-calling-traces)
- [How do I tell whether my judge is biased?](https://tessary.ai/answers/judge-bias-research/how-do-i-tell-whether-my-judge-is-biased)

---

Source: https://tessary.ai/answers/judge-bias-research
All concepts: https://tessary.ai/answers
From Tessary, agent reliability for AI agents in production: https://tessary.ai
