# How do you measure whether a grader itself is accurate?

You run it against a set of cases with a known right answer, human-labeled, and compare the grader's verdict to that label on each one. The result is the same pair of numbers used to evaluate any classifier: precision, how often a case the grader flagged was actually bad, and recall, how much of what was actually bad the grader caught. A grader with low precision wastes a team's time chasing false alarms; one with low recall gives false confidence by staying quiet on real failures.

This is why a labeled golden set earns its keep twice: it tests the agent, and it's the ground truth a grader's own accuracy gets measured against. A grader with no labeled set behind it is an unverified claim about what it catches. Recalibrate the same way you'd recheck any measurement: re-run the grader against the labeled set whenever the agent's behavior, or the grader's own prompt, changes underneath it.

---

Source: https://tessary.ai/answers/graders/how-do-you-measure-whether-a-grader-itself-is-accurate
More on Graders: https://tessary.ai/answers/graders
From Tessary, agent reliability for AI agents in production: https://tessary.ai
