An LLM judge is a model that scores another model's output. Judges have biases that push every verdict the same way, and 2026 research has measured them across enough models to say what matters.
The biggest one is style. "Judging the Judges" (arXiv:2604.23178) tested five judges from four providers and found they favour markdown-formatted answers over plain prose by a wide margin, far more than they favour whichever answer comes first. Length bias depends on the judge: some prefer longer answers, Claude prefers shorter ones, and GPT-4o was neutral. The same study found a mid-tier model with the right debiasing prompt matched or beat frontier judges at about a fifteenth of the cost.
Self-preference, where a judge favours its own model's output, does not go away with capability. A study across 20 models (arXiv:2604.22891) found stronger models were no less biased, and sometimes more. A structured, multi-dimension scoring prompt cut it by about a third.
Two findings change how you should validate a judge. The largest evaluation to date, 21 judges and about 541,000 judgments (arXiv:2606.19544), showed that raw agreement with humans overstates judge quality by 33 to 41 points once you correct for chance, and that judge rankings move by up to 14 places from one benchmark to the next. And on agent tool-calling traces specifically, AgentJudgeBench (arXiv:2608.26623) found every judge tested, small or frontier, landed in the same 77 to 82% agreement band on hard tasks without ground truth. Chain-of-thought and temperature made no difference; a structured rubric helped by up to 6.5 points.
So: measure your judge against your own labelled traces, correct for chance, and expect a ceiling on hard cases that a bigger model won't lift.
8 questions
Answered, plainly.
Two ways to run Tessary.
Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.
Tessary Cloud
We host it for you. Send your first trace with nothing to deploy and no model key.
- traces
- 10,000 per calendar month
- stored trace data
- 1 GB
- retention
- 30 days
- model credit
- $10, one-time, for triage and root-cause analysis
- credit card
- not required
Self-hosted Tessary
Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.
Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md
docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y