# Do different LLM judges agree when scoring the same rubric?

Barely. A 2026 study had 9 judge models score the same 5-dimension rubric against 120 SEO content packs generated from 30 YouTube videos, and found near-zero agreement between the judges, a Krippendorff's alpha of 0.042, close to chance. What it wasn't was random: each judge disagreed with the others in its own stable, repeatable way, consistent enough that a classifier could tell which judge produced a given set of scores with 77% accuracy from the scores alone, rising to 90% once it also saw how the judge phrased its reasoning.

The paper's read is that judges aren't interchangeable readings of one shared standard; each one applies its own implicit theory of what "good" means, and swapping the model behind a rubric changes what that rubric is actually measuring even when its wording never changes. It's a single-author preprint testing content-quality rubrics, not agent task or tool-call grading, so treat it as evidence a rubric alone doesn't guarantee agreement rather than a measured number for this site's own domain. [Running several judges on one trace doesn't average this out either](/answers/llm-as-judge/does-running-more-judges-reduce-error), since correlated judges lose most of the independence that would need.

---

Sources:
- Evaluative Fingerprints: Stable and Systematic Differences in LLM Evaluator Behavior (arXiv:2601.05114): https://arxiv.org/abs/2601.05114 (fetched 2026-09-19)

Source: https://tessary.ai/answers/judge-bias-research/do-different-llm-judges-agree-when-scoring-the-same-rubric
More on Judge bias research: https://tessary.ai/answers/judge-bias-research
From Tessary, agent reliability for AI agents in production: https://tessary.ai
