all answers

Agent reliability

Cause attribution

Cause attribution is the step from "quality dropped" to "this specific change caused it." Detection establishes that a regression happened; attribution names the change responsible. Knowing a regression exists still leaves the whole search for where it came from. That search is what attribution does.

Any regression has a finite set of candidate causes: a code commit, a prompt edit, a model change, a tool whose behavior shifted, or an upstream system the agent depends on. Attribution is an investigation over that set. You gather the failing traces into one cohort, line up the onset of the failure against the change history, and test each plausible candidate against the evidence in the traces until the ones that don't hold are ruled out.

Localization is what makes the search tractable. Failures cluster around the code path that produces them. When the failing traces share one path and the failure started at a known time, the candidate set shrinks from everything shipped recently to the handful of changes that touched that path in that window.

A finished attribution is a grounded explanation: the named cause, the failing traces and findings that support it, and the hypotheses that were checked and ruled out. Because the cause is a specific change in a specific place, the attribution also says where the fix belongs.

8 questions

Answered, plainly.

How do I find out which change caused my agent's quality drop?Pull the traces where quality dropped into one cohort, line the onset up against your change history, and test each candidate until only one explains what the traces show.answer →How do I tell which deploy was live when a session failed?Match the session's timestamp against your deploy or commit log; the version whose window contains that timestamp is the one that produced it.answer →What evidence should a regression investigation include?The named cause, the failing traces and findings that support it, and the hypotheses you checked and ruled out, not just a summary of what changed.answer →What comes back when the cause of a failure cannot be established?Sometimes nothing does. An honest attribution says no change explains the failure rather than naming the closest guess, and it keeps what it ruled out.answer →Why did a downstream agent get worse when its own code never changed?Because the change that broke it happened upstream: another agent or system it depends on changed what it produces, and the failure shows up downstream first.answer →Why doesn't an alert that reports a score moving help me?A score alert says a metric moved, not why. Someone still has to line the failing traces up against commit, prompt, and model history to find the cause.answer →What is cause attribution?Cause attribution is the investigation that names the specific change responsible for a regression, once detection has established that one happened.answer →Can you trust an agent's own account of a failure?No. An agent's account of its own actions is generated text with the same failure modes as any other output, so it can confidently describe a false recovery.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y