Did my agent get worse, or did my judge?

You can’t tell from a falling score alone, so run a fixed, human-labelled anchor set through the current judge on a steady interleave and watch it with its own statistical test, separate from the one watching live traffic. A June 2026 paper found the usual approach, a rolling z-test against a threshold, false-alarmed on 75 percent of streams where nothing had actually changed, which is why a second, independent check on a known-good slice matters.

On a silent provider version bump, the anchor-set method correctly blamed the judge in 60 of 60 runs. On a harder case, a scoring-prompt edit written to look like a real change, it was right in 110 of 120 runs. Re-scoring the anchor costs roughly a fifth to two thirds of grading every item, so it’s a fraction added to the bill, not a second pipeline.

It can’t tell you if the definition of a good answer itself drifted; that still needs a person rereading the anchor labels.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y