If my LLM judge makes mistakes, is my measured pass rate biased?

Yes, unless you correct for it. A judge’s sensitivity, how often it correctly marks a passing output as passing, and specificity, how often it correctly marks a failing one as failing, are never 100 percent, and that error rate biases every pass rate the judge reports. A judge that misses real failures makes an agent look better than it is; one that flags good outputs as failures makes it look worse, and the two errors don’t cancel out just because they point in opposite directions.

A 2025 paper on reporting LLM-judge evaluations fixes this the way a lab corrects an imperfect diagnostic test: measure the judge’s sensitivity and specificity against a small human-labeled calibration set, then use those two numbers to adjust the raw pass rate and put a real confidence interval around it. Reporting a raw pass rate with no correction is reporting the judge’s own error rate mixed in with the agent’s.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y