Can you fix a correlated LLM judge panel?

Yes, if you explicitly model which judges tend to fail on the same cases instead of averaging every vote as if it came from an independent source. A 2026 study tested this against standard weighted majority voting, the usual way to combine six LLM judges’ verdicts, on three real grading tasks: whether a passage answers a question, whether a comment is toxic, and whether a summary is accurate. Weighted majority voting scored 82%, 72%, and 61% on the three tasks; a method that tracks which judges actually agree with each other, not just how accurate each one is alone, and adjusts for the overlap scored 90%, 79%, and 68%, seven to eight points higher on every task.

Nine judges can carry only about two judges’ worth of real signal specifically because judges that share training data or a similar instruction-tuning recipe tend to miss the same cases, and counting votes alone can’t tell that overlap apart from genuine agreement. Modeling the overlap directly recovers most of what naive counting throws away, but it needs the same thing counting does: a labeled set to check the result against.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y