How does an LLM judge compare to human evaluation?

An LLM judge agrees with human graders about as often as two human graders agree with each other: on MT-Bench, GPT-4 matched a human majority verdict 85 percent of the time on non-tie votes, while two human experts agreed with each other only 81 percent of the time on the same data. That’s the paper that established LLM-as-judge as a working method, and its finding was that the judge landed inside the noise floor humans already have with each other, not below it.

That number is an average, not a fact about your traces. A human grader reads a handful of transcripts a day, applying whatever they currently believe counts as good; an LLM judge runs the same rubric on every trace, all day, for a fraction of the cost, the only way to grade at production volume at all. What a same-generation judge can still miss looks like a biased or drifting judge: a preference for its own model family’s phrasing, a flip when the order of two answers swaps, formatting rewarded over substance. The 85 percent figure is a starting point for trusting a judge, not a reason to stop checking it against a human on a sample.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y