How do I pick the confidence threshold for letting a decision run automatically?

There’s no single number for a whole system. TypeSafe’s own guidance is to gate different actions at different levels depending on what a wrong call costs there, a mistake on an FAQ answer and a mistake on a refund don’t deserve the same bar, and to set each threshold by plotting confidence against accuracy on your own labeled data rather than picking a round number like 0.8 in advance.

The pattern is simple once a threshold exists: send the confident cases straight through and route the rest to a person or a more expensive reasoning model, something like if answer.confidence < 0.8: route_to_human_review(ticket). Checking whether that confidence is calibrated in the first place has to happen before any threshold means anything, since a 0.8 that’s actually right 60% of the time isn’t a safe cutoff wherever you set it.

Start conservative and adjust as you observe real results. Tessary’s own classifiers escalate the same way: crossing a threshold writes a finding for triage rather than acting on its own, because the cost of a wrong automatic call is higher than the cost of a short delay.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y