Do different LLM judges agree when scoring the same rubric?

Barely. A 2026 study had 9 judge models score the same 5-dimension rubric against 120 SEO content packs generated from 30 YouTube videos, and found near-zero agreement between the judges, a Krippendorff’s alpha of 0.042, close to chance. What it wasn’t was random: each judge disagreed with the others in its own stable, repeatable way, consistent enough that a classifier could tell which judge produced a given set of scores with 77% accuracy from the scores alone, rising to 90% once it also saw how the judge phrased its reasoning.

The paper’s read is that judges aren’t interchangeable readings of one shared standard; each one applies its own implicit theory of what “good” means, and swapping the model behind a rubric changes what that rubric is actually measuring even when its wording never changes. It’s a single-author preprint testing content-quality rubrics, not agent task or tool-call grading, so treat it as evidence a rubric alone doesn’t guarantee agreement rather than a measured number for this site’s own domain. Running several judges on one trace doesn’t average this out either, since correlated judges lose most of the independence that would need.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y