What happens when a Tessary classifier flags a trace?

It depends on the classifier. Some file a finding on the flag itself, while the rate-tested ones count the flag and file a finding only when a call site’s rate of flags rises above its own normal. Either way, nothing pages anyone yet.

secret_leak is the first kind: one high-confidence leak is a finding. malformed_output, frustration, and groundedness are the second, so one failed schema check, frustrated message, or unsupported answer is counted and nothing more.

LLM judgment enters at the finding, reading it instead of the raw trace. Triage separates a real deviation from legitimate change, the harder half of the job. A tool your team changed on purpose looks identical to drift until something checks it against what actually happened. A high-confidence leak and a frustration finding are filed already ruled and skip triage. A groundedness finding goes through it.

What survives triage becomes a case, and a case is what gets attributed and, if it clears the alerting rules, what reaches a human. A flag that doesn’t survive triage never does.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y