How do I detect agent regressions before users complain?

Grade the content of every production turn against the call site’s own baseline, and alert when the rate shifts rather than on a single bad run. The signal then comes from traces instead of tickets.

Grade what the turn was for: task completion against the end state the user asked for, groundedness on each answer built from retrieved documents, failure rate per tool. A review sample only sees failures that are both common and sampled, so a rare one takes more sessions than a review queue gets through. And one bad run isn’t a regression: the evidence is a rate moving across two windows of comparable traffic, the same before-and-after comparison you run against a change you already suspect.

Tessary runs its classifiers on every trace, writes a finding when a score crosses its threshold, and triage decides whether that finding becomes a case.

The baseline is the limit. A call site with too few comparable turns has nothing to shift against, so it stays quiet until enough traffic has run through it.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y