How do I catch a tool that starts failing more often than it used to?
Track each tool’s failure rate against its own history and alert on a sustained shift, not on any single failed call. An API that starts refusing requests, a dependency that breaks, or a rate limit that begins to bite all show up as a rate change long before every call is failing.
That means a baseline per tool rather than one cutoff applied everywhere, and evidence that accumulates call by call so ordinary noise stays quiet. Tessary’s tool_error classifier works this way: the threshold comes from a false-alarm budget computed for that tool, and a firing names when the shift began and which call sites saw it.
That budget is a design target, not a rate measured on real traffic. It is computed assuming tool failures are independent, and real ones are bursty: one upstream outage fails hundreds of consecutive calls, which inflates the false-alarm rate by an amount the arithmetic cannot price. The measurement against a real corpus has not been run.