How do you measure agent reliability in production?

You measure it as a rate, not a single verdict. Because agents are nondeterministic, one run tells you little: the same input can produce a good answer one time and a bad one the next. What matters is the distribution, so the practical measure is something like the share of production runs that meet the bar you set for correct behavior, tracked over a rolling window so you can see it move. That requires judging the content of each run, not just whether it completed, since a wrong answer and a right one both return a normal status. It also requires a consistent baseline: the same measure applied before and after a change is what lets you tell a real shift in reliability from ordinary run-to-run noise. A single spot check, on one run or one day, cannot do either job.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y