How can an answer be faithful to its context and still be wrong?

Faithfulness only checks the answer against what was retrieved, not against everything the source actually says. An answer built entirely from a partial slice of a document, half a policy, one clause of a contract, one turn of a longer conversation, can accurately represent that slice and still land on the wrong conclusion, because the part that would have changed the answer never made it into context.

That’s a context-presence failure wearing a faithfulness pass. The model didn’t misquote anything; it correctly summarized what it was given, and what it was given wasn’t enough. This is why the three checks have to run separately rather than one score standing in for all of them: a grader that only asks “does the answer match the context” will wave this case through, because by that narrow question, it did.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y