What comes back when the cause of a failure cannot be established?

Sometimes nothing comes back as a named cause, and that’s a legitimate result, not a failed investigation. Not every regression traces to one located change: sometimes the traffic itself shifted, sometimes the signal was measurement noise rather than a real decline, and sometimes the real cause sits outside anything you can see, in a provider’s model update or a dependency you don’t control. A trustworthy investigation reports that honestly instead of naming the closest available guess just to close the case, because a wrong cause sends someone to fix the wrong thing while the real one keeps firing. What you still get is the record of what was checked and ruled out, so the next occurrence doesn’t start from zero, and a note on what would need to be true for a cause to surface, more traces, better logging, a repo connection. No change found is a real, useful answer. A confident wrong one isn’t.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y