How reliable are MCP tool calls in production?

Not very, on average. A 2026 stress test of 100 MCP servers found a median per-call pass rate of 71%, and reliability across servers was bimodal rather than clustered near that median: some were solid, many were not. The number that matters more than the median is what happens when calls chain. At 71% per call, five sequential tool calls succeed end to end only about 18% of the time, and ten calls drop to roughly 3%. An agent does not need one reliable call, it needs every call in its chain to land.

The same study found a separate gap that makes chains worse: 71% of bottom-decile servers have no idempotency protection. Its authors call the most common reliability problem “a successful first call that the agent mis-classified as failed and then re-invoked.” Retrying a call you cannot confirm failed risks duplicating whatever side effect it already had.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y