How long before a call site's baseline is worth anything?

Once each side of the comparison holds roughly 150 comparable turns, the default sample size the duration and cost detectors wait for before calling anything a deviation. There’s no fixed schedule behind that number: it’s a sample-size threshold, not a calendar one. A call site hit on every request clears it in minutes; one on a rare code path can take real wall-clock time, because the count comes from traffic as it actually arrives, never backfilled from history you already had. Baselines are also scoped per call site on purpose: pooling traffic across call sites would let one call site’s normal noise hide another’s real drift, so a low-volume call site waits on its own clock regardless of how busy the rest of your agent is. Silence from a freshly tagged call site almost always means the baseline is still building, not that nothing’s wrong; the sample count is visible while it climbs.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y