How do you know when a grader no longer matches what the agent does?

You know by checking whether the prompt, the tool’s contract, or the output schema the grader was written against has changed, not by watching the pass rate. The pass rate alone won’t tell you, because a grader keeps returning verdicts whether or not the thing it was written to check still exists.

A grader that checks “did the agent ask for an order ID before refunding” is silently wrong the day someone removes the order ID requirement, even though every trace after that keeps passing cleanly. The rate looks healthy because the behavior it’s checking for no longer applies to anyone.

Tie the review to the commit, not the calendar: any change that touches the prompt, tool contract, or output schema a grader depends on is the trigger to reopen it, whether or not its rate has moved yet.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y