Can you fix a correlated LLM judge panel?
Yes, if you explicitly model which judges tend to fail on the same cases instead of averaging every vote as if it came from an independent source. A 2026 study tested this against standard weighted majority voting, the usual way to combine six LLM judges’ verdicts, on three real grading tasks: whether a passage answers a question, whether a comment is toxic, and whether a summary is accurate. Weighted majority voting scored 82%, 72%, and 61% on the three tasks; a method that tracks which judges actually agree with each other, not just how accurate each one is alone, and adjusts for the overlap scored 90%, 79%, and 68%, seven to eight points higher on every task.
Nine judges can carry only about two judges’ worth of real signal specifically because judges that share training data or a similar instruction-tuning recipe tend to miss the same cases, and counting votes alone can’t tell that overlap apart from genuine agreement. Modeling the overlap directly recovers most of what naive counting throws away, but it needs the same thing counting does: a labeled set to check the result against.