# Does a judge panel's verdict depend on which model families sit on it?

Yes. A study built panels from four open-weight model families, Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5, and ran 9,312 pairwise judgments to isolate what a judge's own family does to a verdict, separate from how good an answer actually is. Every family showed the same pattern, a 3.4 to 8.4 percentage point edge for answers from their own family, and relative to a family-balanced reference panel, swapping which families sat on the panel changed 18.5% of its pairwise verdicts.

The bias survives the obvious fixes. It held up under panel-based quality controls, an independent human check, and a higher-precision, unquantized replication, and it isn't fully explained by judges just preferring familiar-sounding text: correcting for that only accounts for 61% of it. [Running more judges on the same trace doesn't fix correlated errors either](/answers/llm-as-judge/does-running-more-judges-reduce-error), and here the correlation runs along family lines specifically. The result is on open-weight models only; frontier closed models weren't tested.

---

Sources:
- "Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels" (arXiv:2609.17857): https://arxiv.org/abs/2609.17857 (fetched 2026-09-25)

Source: https://tessary.ai/answers/judge-bias-research/does-judge-panel-composition-change-the-verdict
More on Judge bias research: https://tessary.ai/answers/judge-bias-research
From Tessary, agent reliability for AI agents in production: https://tessary.ai
