What is Tessary's groundedness classifier?

It’s a built-in classifier that checks your agent’s answers against the documents it retrieved, or against the prompt when it retrieved none, and flags the sentences that source doesn’t support.

A sentence counts as unsupported when the source contradicts it or never says it at all. The scorer is an open model, tessaryai/groundedness-classifier-v1, which reads the source and the whole answer together in one pass and scores each sentence for how likely it is to be unsupported. It isn’t an LLM judge: it runs on a GPU you provide, a Mac for development or an AWS instance for production, with no provider key and no per-answer charge.

On the RAGTruth benchmark, its model card reports that at 0.975, the threshold Tessary flags at, it catches 39% of unsupported answers, and 83% of the answers it flags are unsupported. So one flag isn’t a finding. Each call site learns its normal share of traces with a flagged answer, and a finding opens when that share rises, then goes to triage.

The limit is recall: at that threshold it misses most unsupported answers, so a quiet call site doesn’t mean every answer on it was grounded.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y