How do I check whether a model's confidence is calibrated on my own traffic?
Label a sample of your real inputs with the right answer, run the model on them, then bucket the predictions by confidence and compare each bucket’s average confidence to how often it was actually right. That comparison is a reliability diagram, the tool Guo et al. (2017) used to show modern neural networks are poorly calibrated. Its one-number summary, expected calibration error, is the weighted average of the gaps. When one bad bucket is what hurts you, maximum calibration error, the worst single gap, is the better number.
It’s the same labeled-set work as measuring a grader’s accuracy, and TypeSafe’s build guide tells Jev users to test thresholds by plotting confidence against accuracy on their own data.
The diagram doesn’t show how many predictions sit in each bucket, so print the counts beside it. A threshold doesn’t carry between question shapes: TypeSafe’s jaggedness notes show one refund question asked as a yes-or-no probability returning 0.22 and as a two-option choice returning 0.01 for yes. And an alias moves when a new version ships, so TypeSafe’s models page says to pin the version your thresholds were tuned against.