Does a judge panel's verdict depend on which model families sit on it?
Yes. A study built panels from four open-weight model families, Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5, and ran 9,312 pairwise judgments to isolate what a judge’s own family does to a verdict, separate from how good an answer actually is. Every family showed the same pattern, a 3.4 to 8.4 percentage point edge for answers from their own family, and relative to a family-balanced reference panel, swapping which families sat on the panel changed 18.5% of its pairwise verdicts.
The bias survives the obvious fixes. It held up under panel-based quality controls, an independent human check, and a higher-precision, unquantized replication, and it isn’t fully explained by judges just preferring familiar-sounding text: correcting for that only accounts for 61% of it. Running more judges on the same trace doesn’t fix correlated errors either, and here the correlation runs along family lines specifically. The result is on open-weight models only; frontier closed models weren’t tested.