How do I detect agent regressions before users complain?
Grade the content of every production turn against the call site’s own baseline, and alert when the rate shifts rather than on a single bad run. The signal then comes from traces instead of tickets.
Grade what the turn was for: task completion against the end state the user asked for, groundedness on each answer built from retrieved documents, failure rate per tool. A review sample only sees failures that are both common and sampled, so a rare one takes more sessions than a review queue gets through. And one bad run isn’t a regression: the evidence is a rate moving across two windows of comparable traffic, the same before-and-after comparison you run against a change you already suspect.
Tessary runs its classifiers on every trace, writes a finding when a score crosses its threshold, and triage decides whether that finding becomes a case.
The baseline is the limit. A call site with too few comparable turns has nothing to shift against, so it stays quiet until enough traffic has run through it.