Does a higher pass rate mean an eval suite has better coverage?
No. Pass rate is computed only over the cases already in the suite, so adding cases the agent already handles raises the score and leaves coverage exactly where it was, or even widens the gap if those cases crowd out ones that would have caught a real failure. Coverage is a property of how well the suite’s inputs match what production actually sends, not of the score, and the two move independently: a suite can raise its pass rate every quarter while the input types it never tests grow just as fast.
The fix isn’t more cases, it’s checking what’s missing. Sample recent production traces, cluster them by intent, and compare the clusters against the cases you already have; a cluster with no case pointing at it is a hole no amount of passing tests will show you. That’s a different question from what it costs to grade what you already have, and cheaper than assuming a high score means the gap already closed.