Does tool_error flag a single failed tool call?

No. tool_error watches a tool’s failure rate over time, not any individual call, so one failure sitting inside a tool’s normal error rate doesn’t move the statistic far enough to fire.

Most tools fail sometimes: a timeout, a malformed request, a flaky dependency. Treating every one of those as an incident would bury a real problem under noise from calls that were never going to cross any bar. Evidence instead accumulates call by call against each tool’s own baseline, and it’s a rate that’s shifted, several failures clustering where the tool used to be reliable, that crosses the threshold and becomes a finding.

That also means it won’t catch a single high-stakes failure the moment it happens; a lone bad call inside an otherwise healthy rate is invisible to this classifier by design. Catching one specific failure regardless of the surrounding rate is a different kind of check.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y