all answers

Agent reliability

Regression detection

A regression is a drop in agent quality caused by a change. Regression detection is the practice of noticing that drop by comparing the agent's outcome after a change against its outcome before.

Two properties of agents make this harder than in conventional software. First, agent output is non- deterministic, so a single bad run carries little information. The signal is distributional: a failure rate or quality score that shifts after a change ships. Second, the change that causes a regression can come from anywhere in the system around the model. A refactor that changes what context gets assembled, a dependency bump that alters tool output formatting, or a config edit that changes retry behavior can each regress the agent's behavior.

There's also a judgment problem the statistics alone can't settle. Teams change their agents on purpose, and an intended improvement moves the same numbers a defect does. Whether a detected shift counts as a regression depends on what the agent was meant to do after the change.

9 questions

Answered, plainly.

Can a production regression happen without a code change?Yes. The model behind a stable API name can change on the provider's side, and a dependency bump or config edit can shift behavior with no prompt touched.answer →Does a passing unit test rule out a regression?No. A unit test asserts your code executed the right path, such as calling a tool with valid arguments; it says nothing about whether the model's behavior stayed good.answer →How do you detect an agent regression after it's already in production?Compare a window of production runs against a reference window from before the suspected change, on the same quality measure, and watch the gap hold up over enough runs.answer →Is every quality shift after a change a regression?No. A regression is an unwanted drop; the same movement in a metric is expected, even desired, when the change was made on purpose to produce it.answer →What regressions will a CI gate never catch?Any regression whose cause never shows up in a diff: a model updated behind a stable API, a dependency's behavior shifting, or a downstream service changing its output.answer →Why isn't a single bad run enough to call it a regression?Agent output is nondeterministic, so one bad run can be ordinary variance. A regression is a shift in the rate of bad runs, visible only across many of them.answer →How do I detect agent regressions before users complain?Grade every production turn against the call site's own baseline and alert on a shift in the rate, not on one bad run, so the evidence arrives from traces rather than tickets.answer →What is regression testing for an AI agent?Re-running a fixed set of cases, each with an expected behavior, after every change, to catch behavior that used to work and stopped.answer →What is drift in an AI agent?Any change in an agent's behavior that nobody deliberately made, whether from a model update, a shifting user base, or the world it describes going stale.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y