Does a bigger model make a better judge?

No. A study measuring self-preference bias across 20 mainstream models found capability uncorrelated with bias, and in some cases negatively correlated, meaning a stronger judge model can favor its own outputs more than a weaker one does, not less. What actually cut the bias, by about 31.5% on average, was a structured, multi-dimension scoring prompt, a change to how the judge is asked to score, not a bigger model doing the scoring.

Scale doesn’t buy stability either. A separate study spanning 21 judges from nine providers and roughly 541,000 judgments found judge rankings shift by up to 14 places depending on which benchmark measures them, and that raw agreement with human labels overstates judge quality by 33 to 41 percentage points once corrected for chance agreement. Neither finding tracks model size. The fix in both studies is the same: calibrate the judge against your own labeled cases instead of trusting a bigger name on the model card, and structure the scoring prompt rather than just upgrading it.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y