How do you measure whether a grader itself is accurate?
You run it against a set of cases with a known right answer, human-labeled, and compare the grader’s verdict to that label on each one. The result is the same pair of numbers used to evaluate any classifier: precision, how often a case the grader flagged was actually bad, and recall, how much of what was actually bad the grader caught. A grader with low precision wastes a team’s time chasing false alarms; one with low recall gives false confidence by staying quiet on real failures.
This is why a labeled golden set earns its keep twice: it tests the agent, and it’s the ground truth a grader’s own accuracy gets measured against. A grader with no labeled set behind it is an unverified claim about what it catches. Recalibrate the same way you’d recheck any measurement: re-run the grader against the labeled set whenever the agent’s behavior, or the grader’s own prompt, changes underneath it.