Does groundedness check every LLM call, or just the final answer?

Neither: it checks the output of every model call at a call site whose shape is rag_answer, summarize, or extract, wherever that call sits in the trace, and no other call.

Those three shapes are the calls that answer from source material: retrieved documents, or a document handed over in the prompt. A planning step, a call that picks the next tool, or an open-ended draft has no source to be ungrounded from, so it’s never scored. A call with no call site isn’t checked either. If an agent answers from documents and then rewrites that answer for the user, the rewrite is checked only if its call site has one of those shapes too.

Inside a checked call, the whole answer is read against its source in one pass. An answer that only greets, thanks, or asks the user something is skipped, since there’s nothing in it to check.

An invented detail from an unchecked step can still be flagged if it reaches a checked answer built from retrieved documents, because that answer is read against the documents, not against what the earlier step wrote.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y