all answers

Tessary

Tessary classifiers

Tessary's classifiers are the cheap checks that run on every production trace. Sampling misses low-probability issues, and in a mature agent nearly every remaining issue is low-probability. So finding them means checking everything, which is only affordable when one check costs almost nothing.

Each trace is scored by the built-ins: frustration, groundedness, secret_leak, malformed_output, behavior_drift, duration_drift, cost_drift, tool_error, and sop_conformance. When a score crosses its confidence threshold, a finding is created for further analysis; LLM judgment comes in only at that escalation step, never by default. Custom classifiers, fine-tuned for your agent's own failure modes, run alongside them.

Low-probability issues surface as they happen instead of weeks later through user reports, which is what keeps your critical agents reliable.

8 questions

Answered, plainly.

How much does it cost to run a classifier on every trace?Nothing, on the self-hosted platform. You pay your model provider for triage, RCA, and frustration scoring, and pay for your own GPU time if you run groundedness.answer →What happens when a Tessary classifier flags a trace?It depends on the classifier. Some file a finding on the flag itself, rate-tested ones only when a call site's flag rate rises, and a finding pages no one on its own.answer →What is a Tessary classifier?A narrow, cheap check that scores one production trace for a single property, cheap enough to run on all of it instead of a sample.answer →Why do Tessary's classifiers produce false positives?Because each classifier scores one narrow property of a trace and fires on a threshold. A flag is not a verdict: a rising rate, triage, or both decide what is real.answer →Why does Tessary check every trace instead of sampling?Sampling measures how often something happens; it is a poor tool for catching something rare, and in a mature agent nearly everything left to catch is rare.answer →Does Tessary truncate a long trace before grading it?Partly. Frustration clips long messages and groundedness shortens documents to fit one pass, while triage and RCA include fewer items rather than clip any one.answer →Do I need my own LLM provider key for Tessary to grade traces?Self-hosted, yes for triage and RCA: they run on your org's own model provider key. Groundedness needs no key, only a model you run yourself.answer →What's the difference between an LLM and a classifier?An LLM reasons over open text and answers almost anything; a classifier is trained to score one narrow property, cheaply enough to run on every trace.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y