Does running more LLM judges on the same trace reduce error?
Not by much. A 2026 study called “Nine Judges, Two Effective Votes” found a panel of nine frontier LLM judges carried only about two judges’ worth of independent signal, since models trained similarly tend to fail on the same items. The panel’s accuracy landed 8 to 22 percentage points below nine truly independent votes, and the single best judge matched or beat the full panel.
The bottleneck is correlation between judges, not the aggregation method: models trained on overlapping data and instruction-tuned in similar ways tend to fail on the same inputs. Measuring a grader’s own accuracy still needs a labeled set to check against; stacking judges without that check just multiplies the cost of being wrong together.