If my LLM judge makes mistakes, is my measured pass rate biased?
Yes, unless you correct for it. A judge’s sensitivity, how often it correctly marks a passing output as passing, and specificity, how often it correctly marks a failing one as failing, are never 100 percent, and that error rate biases every pass rate the judge reports. A judge that misses real failures makes an agent look better than it is; one that flags good outputs as failures makes it look worse, and the two errors don’t cancel out just because they point in opposite directions.
A 2025 paper on reporting LLM-judge evaluations fixes this the way a lab corrects an imperfect diagnostic test: measure the judge’s sensitivity and specificity against a small human-labeled calibration set, then use those two numbers to adjust the raw pass rate and put a real confidence interval around it. Reporting a raw pass rate with no correction is reporting the judge’s own error rate mixed in with the agent’s.