# Does running more LLM judges on the same trace reduce error?

Not by much. A 2026 study called "Nine Judges, Two Effective Votes" found a panel of nine frontier LLM judges carried only about two judges' worth of independent signal, since models trained similarly tend to fail on the same items. The panel's accuracy landed 8 to 22 percentage points below nine truly independent votes, and the single best judge matched or beat the full panel.

The bottleneck is correlation between judges, not the aggregation method: models trained on overlapping data and instruction-tuned in similar ways tend to fail on the same inputs. [Measuring a grader's own accuracy](/answers/graders/how-do-you-measure-whether-a-grader-itself-is-accurate) still needs a labeled set to check against; stacking judges without that check just multiplies the cost of being wrong together.

---

Sources:
- "Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels" (arXiv:2605.29800): https://arxiv.org/abs/2605.29800 (fetched 2026-09-09)

Source: https://tessary.ai/answers/llm-as-judge/does-running-more-judges-reduce-error
More on LLM as judge: https://tessary.ai/answers/llm-as-judge
From Tessary, agent reliability for AI agents in production: https://tessary.ai
