Does duration_drift use a fixed latency threshold?

No. duration_drift has no absolute threshold anywhere in it, like flagging anything slower than five seconds. A single-tool lookup and a thirty-step research run can both be completely normal while sitting two orders of magnitude apart, so a fixed cutoff would either miss the agent that’s always been slow or fire constantly on the one that’s supposed to be fast.

Instead, each call site and each tool is compared only against its own recent history. What counts as drift is a real shift away from what that specific bucket normally does, reported as a multiplicative change like 1.4x slower rather than a raw number of seconds. That’s also why a brand new call site takes a while to be useful: there’s no history to compare against until it’s built up enough of its own traffic to know what normal looks like for it.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y