What's the difference between an LLM and a classifier?

An LLM reasons over open-ended text and can answer almost anything you ask it; a classifier is trained to score one specific, narrow property and nothing else, which is what makes it cheap enough to run on every trace. Tessary’s own built-ins show the split: groundedness runs a small open model trained only to score whether a sentence is supported by its retrieved documents, and frustration runs TypeSafe’s Jev, a decision model trained to answer one typed question rather than hold a conversation. Both cost a fraction of an LLM call, because neither is doing general reasoning; each is scoring the one thing it was built to score. An LLM judge earns its higher cost back when a check genuinely needs judgment a narrow classifier can’t give: tone, whether an answer actually addresses what was asked, a policy with exceptions. That’s why Tessary reaches for one only at the escalation step, after a cheap classifier has already flagged something worth a closer look.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y