What does a behavior drift baseline catch, and what does it miss?

It catches a new or reordered step almost every time, and misses a real share of the times an agent quietly stops doing something it used to do.

The reason is structural, not a tuning problem. An agent calling a tool it’s never used, or taking a path it’s never taken, leaves something new in the trace for the baseline to notice. An agent dropping a step it always ran leaves nothing: the trace just looks like a shorter version of normal, and normal is exactly what a behavior baseline is built to let through. New actions, swapped steps, and reordered ones are the strong catches. A quiet omission is the classifier’s honest weak spot.

If a step matters for correctness and not just for shape, pairing this with a check that asserts the step happened covers the gap drift detection alone won’t.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y