all answers

Tessary

Tessary groundedness

Groundedness is one of Tessary's built-in classifiers. For each answer built from retrieved documents, it checks whether what the agent said is supported by those documents, sentence by sentence.

It catches the shape most production hallucinations take: an answer that asserts things its documents don't support while reading as fluent and confident. A sentence counts as unsupported when the documents contradict it or never say it at all.

The classifier is a small open model, not an LLM judge. It reads the retrieved documents and the whole answer together in one pass, so a claim that draws on two documents is judged against both. It scores each sentence for how likely it is to be unsupported, and an answer is flagged when any sentence crosses a high bar.

It judges the answer against the documents, not the documents against the world. If retrieval returns a stale or wrong document, an answer that faithfully repeats it reads as grounded. An answer that is true but came from the model's own knowledge can still be flagged.

7 questions

Answered, plainly.

Can a factually correct answer still get flagged as ungrounded?Yes. Groundedness checks whether an answer is supported by its documents, not whether the answer is true, and those are different questions.answer →Does groundedness check every LLM call, or just the final answer?Neither. It checks every model call at a call site shaped rag_answer, summarize, or extract, wherever it sits in the trace. Planning and tool-choice calls aren't checked.answer →Does the groundedness classifier check whether the context itself is correct?No. It checks whether the answer is supported by the documents it was given, not whether those documents are accurate, so a wrong or stale source can still score grounded.answer →Does the groundedness classifier only work on agents that use RAG?No. It checks RAG answers against their retrieved documents, and summarize or extract calls against their prompt. RAG answers built only from tool output aren't checked.answer →What is Tessary's groundedness classifier?A built-in classifier that checks your agent's answers against the documents it retrieved, or the prompt when it retrieved none, and flags the sentences that source doesn't support.answer →How long before Tessary's groundedness classifier starts flagging answers?It starts once its reference holds 100 traces, down from 200; it already kept learning until 1,000 before this change, and still does.answer →Can Tessary's groundedness classifier check a very long answer?Up to a point. It reads the answer and its documents in one pass capped at 8,192 tokens; an answer long enough to leave no room for a document isn't scored at all.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y