What is Tessary's tool_error classifier?
tool_error watches the rate at which a tool’s calls fail, not any single call. A call counts as failed only on structural evidence: an error status on the span, a recorded exception, or a result that declares itself an error. Output text that merely mentions the word error never counts.
A detection is one finding per tool, and it carries the tool’s name, when the rate moved, which failure patterns drove it, and which call sites the traffic came through. From there it goes the way any classifier’s finding does.
Two limits. Because the evidence is structural, a tool that returns success with a wrong result is invisible to this check, which is a different failure needing a different one. And the per-tool alert threshold comes from a false-alarm budget that has never been measured against real traffic.