# Does a more capable LLM judge show less self-preference bias?

No. A study measuring self-preference bias across 20 mainstream LLM judges found that a model's capability and its bias were uncorrelated at best, and in some cases moved together instead of apart: the researchers describe advanced capability as "often uncorrelated, or even negatively correlated," with lower bias. Swapping in a newer, stronger judge is not a fix on its own.

What did cut the bias was the prompt, not the model. Splitting a verdict into a structured, multi-dimension format, separate judgments instead of one overall score, reduced measured self-preference bias by 31.5% on average across the judges tested, more than any capability difference between them.

That's a different failure than [a judge favoring text that merely sounds like its own model family](/answers/llm-as-judge/does-an-llm-judge-favor-its-own-models-outputs): this one tracks how a verdict is prompted for, not how familiar the writing sounds. Restructure the rubric before reaching for a bigger judge.

---

Sources:
- "Quantifying and Mitigating Self-Preference Bias of LLM Judges" (arXiv:2604.22891): https://arxiv.org/abs/2604.22891 (fetched 2026-09-16)

Source: https://tessary.ai/answers/judge-bias-research/does-a-more-capable-llm-judge-show-less-self-preference-bias
More on Judge bias research: https://tessary.ai/answers/judge-bias-research
From Tessary, agent reliability for AI agents in production: https://tessary.ai
