Does checking whether an answer is grounded require the model that generated it?
No. A paper published this month trained a detector, called a Grounding Probe, on the hidden states of a model that never generated the answer at all, an “observer” that only reads the context, the question, and someone else’s response. Across four different observer models, it reached 0.88 to 0.89 AUROC on RAGTruth, a standard test set for this kind of error, and adding a supervised detector alongside it pushed that to 0.92. One probe held up across six different generators, including ones it never saw in training.
That removes a real constraint. Every earlier hidden-state method read the generating model’s own activations, so a closed-weight generator, or a generator that changes, broke the detector. An observer model sidesteps that the same way Tessary’s groundedness classifier already does, by scoring the answer against its source with a separate model rather than interrogating the one that wrote it.
Asking the observer model outright, instead of reading its hidden state, cost at least 0.166 AUROC in every model tested. The hidden state carries more than the model says when asked directly.