# If my LLM judge makes mistakes, is my measured pass rate biased?

Yes, unless you correct for it. A judge's sensitivity, how often it correctly marks a passing output as passing, and specificity, how often it correctly marks a failing one as failing, are never 100 percent, and that error rate biases every pass rate the judge reports. A judge that misses real failures makes an agent look better than it is; one that flags good outputs as failures makes it look worse, and the two errors don't cancel out just because they point in opposite directions.

A 2025 paper on reporting LLM-judge evaluations fixes this the way a lab corrects an imperfect diagnostic test: measure the judge's sensitivity and specificity against a small human-labeled calibration set, then use those two numbers to adjust the raw pass rate and put a real confidence interval around it. Reporting a raw pass rate with no correction is reporting the judge's own error rate mixed in with the agent's.

---

Sources:
- Lee & Zeng, "How to Correctly Report LLM-as-a-Judge Evaluations" (arXiv:2511.21140): https://arxiv.org/abs/2511.21140 (fetched 2026-09-04)

Source: https://tessary.ai/answers/llm-as-judge/does-judge-error-rate-bias-my-measured-pass-rate
More on LLM as judge: https://tessary.ai/answers/llm-as-judge
From Tessary, agent reliability for AI agents in production: https://tessary.ai
