Does Tessary use an LLM to check for malformed output?

No. Checking a structured output against its declared schema is deterministic: either it parses and matches the shape, or it doesn’t, so there’s nothing left for a model to judge. No LLM call happens.

That matters for two reasons. Cost is the first: a deterministic check runs at negligible per-trace cost, which is what lets Tessary validate every trace instead of a sample, unlike the model judgment reserved for traces a classifier has already flagged. Confidence is the second: a judged check can be uncertain about how bad something is, but a schema either matches or it doesn’t, so every malformed_output finding carries the actual mismatch as evidence rather than a probability.

The tradeoff is scope. The classifier can tell you an output broke its contract; it can’t tell you whether the content inside a technically valid output is any good. That’s a different, harder question, and not this classifier’s job.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y