Why does Tessary check every trace instead of sampling?

Because sampling is built to measure a rate, and by the time an agent is mature, nearly everything left to catch is rare, exactly what a sample misses. Tessary checks every trace because the checks are cheap enough to afford it.

Sampling is a rate instrument: accurate and cheap for telling you how often something happens, nearly useless for catching one instance of something uncommon. The failures a team already fixed are the frequent ones, so what’s left as an agent matures is low-probability by construction.

What matters is what one check has to cost before reading everything is affordable. Deep judgment on every trace isn’t that. The classifiers that run by default are rule checks and statistical tests over stored spans. Groundedness, which you switch on, runs a public model on a GPU you provide rather than calling an LLM per answer. It still reads only the answers it can check: those on call sites that answer from retrieved documents, summarize, or extract.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y