Groundedness is one of Tessary's built-in classifiers. For each answer built from retrieved documents, it checks whether what the agent said is supported by those documents, sentence by sentence.
It catches the shape most production hallucinations take: an answer that asserts things its documents don't support while reading as fluent and confident. A sentence counts as unsupported when the documents contradict it or never say it at all.
The classifier is a small open model, not an LLM judge. It reads the retrieved documents and the whole answer together in one pass, so a claim that draws on two documents is judged against both. It scores each sentence for how likely it is to be unsupported, and an answer is flagged when any sentence crosses a high bar.
It judges the answer against the documents, not the documents against the world. If retrieval returns a stale or wrong document, an answer that faithfully repeats it reads as grounded. An answer that is true but came from the model's own knowledge can still be flagged.
7 questions
Answered, plainly.
Two ways to run Tessary.
Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.
Tessary Cloud
We host it for you. Send your first trace with nothing to deploy and no model key.
- traces
- 10,000 per calendar month
- stored trace data
- 1 GB
- retention
- 30 days
- model credit
- $10, one-time, for triage and root-cause analysis
- credit card
- not required
Self-hosted Tessary
Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.
Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md
docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y