How can an answer be faithful to its context and still be wrong?

Faithfulness only checks the answer against what was retrieved, not against everything the source actually says. An answer built entirely from a partial slice of a document, half a policy, one clause of a contract, one turn of a longer conversation, can accurately represent that slice and still land on the wrong conclusion, because the part that would have changed the answer never made it into context.

That’s a context-presence failure wearing a faithfulness pass. The model didn’t misquote anything; it correctly summarized what it was given, and what it was given wasn’t enough. This is why the three checks have to run separately rather than one score standing in for all of them: a grader that only asks “does the answer match the context” will wave this case through, because by that narrow question, it did.

keep reading

More on this.

Send us the traces you already emit.