Did my agent get worse, or did my judge?

You can’t tell from a falling score alone, so run a fixed, human-labelled anchor set through the current judge on a steady interleave and watch it with its own statistical test, separate from the one watching live traffic. A June 2026 paper found the usual approach, a rolling z-test against a threshold, false-alarmed on 75 percent of streams where nothing had actually changed, which is why a second, independent check on a known-good slice matters.

On a silent provider version bump, the anchor-set method correctly blamed the judge in 60 of 60 runs. On a harder case, a scoring-prompt edit written to look like a real change, it was right in 110 of 120 runs. Re-scoring the anchor costs roughly a fifth to two thirds of grading every item, so it’s a fraction added to the bill, not a second pipeline.

It can’t tell you if the definition of a good answer itself drifted; that still needs a person rereading the anchor labels.

sources

keep reading

More on this.

Send us the traces you already emit.