How do I catch agent hallucinations when my error rate stays flat?
Check the content of every LLM call in production against its retrieved context, because a hallucination returns a clean status and only a content check can see it.
Which content check to run is the easier question. The harder one is where it runs. A hallucination depends on what was retrieved for that specific request, so a fixed test set exercises retrievals your users never trigger, and a green suite says nothing about the answers going out right now.
Coverage is the other half. Hallucinations are rare per call, so a check on a thin manual sample can run for weeks without landing on one. That puts a cost ceiling on the check itself: it has to be cheap enough to run on all of production rather than a slice of it.
The error-rate dashboard has no part in this. It reads whether the run completed, not what it said.