Does calibration hold when my production traffic drifts from what the model was trained on?
No, not automatically. A calibration guarantee only covers the traffic mix it was measured on, and an independent test of Jev’s certified thresholds, jev-certify, shows how far it can miss once that mix moves. Calibrated on a support taxonomy where 13.0% of traffic fell outside the known categories, a scope-detection gate built to hold a 5% error rate missed by 3.6x, 16.57% measured, once real traffic ran 42.9% outside the taxonomy. A second deployment held its certified 1.75% error rate cleanly at 20 known intents, then broke outright, a 100% error rate, the moment real traffic arrived carrying the other 130 intents the calibration set never saw.
The project’s own conclusion is that a certificate says nothing about traffic it wasn’t calibrated on, and that a drift monitor to catch the moment production traffic moves away from it hasn’t been built yet. The same gap shows up in any agent that changes without a deploy on your side: recalibrate against your current traffic mix rather than a number measured months ago, and expect the gap to widen exactly when that mix shifts most.