Do different LLM judges agree when scoring the same rubric?
Barely. A 2026 study had 9 judge models score the same 5-dimension rubric against 120 SEO content packs generated from 30 YouTube videos, and found near-zero agreement between the judges, a Krippendorff’s alpha of 0.042, close to chance. What it wasn’t was random: each judge disagreed with the others in its own stable, repeatable way, consistent enough that a classifier could tell which judge produced a given set of scores with 77% accuracy from the scores alone, rising to 90% once it also saw how the judge phrased its reasoning.
The paper’s read is that judges aren’t interchangeable readings of one shared standard; each one applies its own implicit theory of what “good” means, and swapping the model behind a rubric changes what that rubric is actually measuring even when its wording never changes. It’s a single-author preprint testing content-quality rubrics, not agent task or tool-call grading, so treat it as evidence a rubric alone doesn’t guarantee agreement rather than a measured number for this site’s own domain. Running several judges on one trace doesn’t average this out either, since correlated judges lose most of the independence that would need.