# Can a brand-new eval suite already have a coverage gap?

Yes. A coverage gap is the distance between the inputs an eval set contains and the inputs production actually sends, and that distance can exist the moment the set is written, before anything has gone stale. [Staleness](/answers/eval-datasets/how-do-i-tell-an-eval-case-has-gone-stale) is a set that used to match production and drifted away from it; a coverage gap is a mismatch on day one, because the cases were built from what someone expected the agent to face rather than sampled from traffic it actually handles.

A high pass rate doesn't tell you the gap isn't there. Hamel Husain and Shreya Shankar's AI evals FAQ warns that passing 100 percent of your evals usually means the suite isn't challenging the system, and cases written from imagination can't tell you how often a failure actually occurs in production, only that someone thought it might. A rate computed entirely inside the gap looks the same as a rate computed outside it.

The fix isn't writing more cases from a whiteboard. [A score is a claim about the eval set it ran on](/answers/graders/what-an-eval-score-actually-claims), so closing the gap means pulling cases from real traffic, not padding the imagined ones.

---

Sources:
- Hamel Husain and Shreya Shankar, "AI Evals FAQ": https://hamel.dev/blog/posts/evals-faq/ (fetched 2026-09-09)

Source: https://tessary.ai/answers/eval-datasets/can-a-brand-new-eval-suite-already-have-a-coverage-gap
More on Eval datasets: https://tessary.ai/answers/eval-datasets
From Tessary, agent reliability for AI agents in production: https://tessary.ai
