Does calibration hold when my production traffic drifts from what the model was trained on?

No, not automatically. A calibration guarantee only covers the traffic mix it was measured on, and an independent test of Jev’s certified thresholds, jev-certify, shows how far it can miss once that mix moves. Calibrated on a support taxonomy where 13.0% of traffic fell outside the known categories, a scope-detection gate built to hold a 5% error rate missed by 3.6x, 16.57% measured, once real traffic ran 42.9% outside the taxonomy. A second deployment held its certified 1.75% error rate cleanly at 20 known intents, then broke outright, a 100% error rate, the moment real traffic arrived carrying the other 130 intents the calibration set never saw.

The project’s own conclusion is that a certificate says nothing about traffic it wasn’t calibrated on, and that a drift monitor to catch the moment production traffic moves away from it hasn’t been built yet. The same gap shows up in any agent that changes without a deploy on your side: recalibrate against your current traffic mix rather than a number measured months ago, and expect the gap to widen exactly when that mix shifts most.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y