What causes a retry storm in an agent?

A retry storm happens when the agent’s retry policy treats a permanent error as if it were transient, so it keeps retrying a call that can never succeed. The signature is a tool call repeated with near-identical arguments, each attempt failing the same way: an unknown tool name, an argument that fails schema validation, a 400 rather than a 429 or a timeout. Retry counts above one on that kind of non-transient error are the tell.

The most common trigger is a hallucinated tool name. In a 200-task ReAct benchmark measured by Towards Data Science, 466 of 513 retries, 90.8 percent of the retry budget, hit a tool the model had invented. A retry policy built for flaky networks and rate limits will run that call forever, because from its point of view every attempt looks like an ordinary failure worth trying again.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y