How do I catch a tool that starts failing more often than it used to?

Track each tool’s failure rate against its own history and alert on a sustained shift, not on any single failed call. An API that starts refusing requests, a dependency that breaks, or a rate limit that begins to bite all show up as a rate change long before every call is failing.

That means a baseline per tool rather than one cutoff applied everywhere, and evidence that accumulates call by call so ordinary noise stays quiet. Tessary’s tool_error classifier works this way: the threshold comes from a false-alarm budget computed for that tool, and a firing names when the shift began and which call sites saw it.

That budget is a design target, not a rate measured on real traffic. It is computed assuming tool failures are independent, and real ones are bursty: one upstream outage fails hundreds of consecutive calls, which inflates the false-alarm rate by an amount the arithmetic cannot price. The measurement against a real corpus has not been run.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y