# How does an LLM judge compare to human evaluation?

An LLM judge agrees with human graders about as often as two human graders agree with each other: on MT-Bench, GPT-4 matched a human majority verdict 85 percent of the time on non-tie votes, while two human experts agreed with each other only 81 percent of the time on the same data. That's the paper that established LLM-as-judge as a working method, and its finding was that the judge landed inside the noise floor humans already have with each other, not below it.

That number is an average, not a fact about your traces. A human grader reads a handful of transcripts a day, applying whatever they currently believe counts as good; an LLM judge runs the same rubric on every trace, all day, for a fraction of the cost, the only way to grade at production volume at all. What a same-generation judge can still miss looks like [a biased or drifting judge](/answers/judge-bias-research/how-do-i-tell-whether-my-judge-is-biased): a preference for its own model family's phrasing, a flip when the order of two answers swaps, formatting rewarded over substance. The 85 percent figure is a starting point for trusting a judge, not a reason to stop checking it against a human on a sample.

---

Sources:
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv:2306.05685): https://arxiv.org/abs/2306.05685 (fetched 2026-09-26)

Source: https://tessary.ai/answers/llm-as-judge/how-does-an-llm-judge-compare-to-human-evaluation
More on LLM as judge: https://tessary.ai/answers/llm-as-judge
From Tessary, agent reliability for AI agents in production: https://tessary.ai
