How does an LLM judge compare to human evaluation?
An LLM judge agrees with human graders about as often as two human graders agree with each other: on MT-Bench, GPT-4 matched a human majority verdict 85 percent of the time on non-tie votes, while two human experts agreed with each other only 81 percent of the time on the same data. That’s the paper that established LLM-as-judge as a working method, and its finding was that the judge landed inside the noise floor humans already have with each other, not below it.
That number is an average, not a fact about your traces. A human grader reads a handful of transcripts a day, applying whatever they currently believe counts as good; an LLM judge runs the same rubric on every trace, all day, for a fraction of the cost, the only way to grade at production volume at all. What a same-generation judge can still miss looks like a biased or drifting judge: a preference for its own model family’s phrasing, a flip when the order of two answers swaps, formatting rewarded over substance. The 85 percent figure is a starting point for trusting a judge, not a reason to stop checking it against a human on a sample.