Why does an AI agent's success rate drop so much on a task that takes longer?

Because each extra step or minute of work carries roughly the same chance of failure, and those chances compound the way radioactive decay does. A 2025 analysis of agent benchmark data found that treating each minute a task would take a human as carrying a constant hazard of failure predicts real success rates surprisingly well: an agent that completes a one-hour task with 50 percent probability drops to about 25 percent on a two-hour task and about 6 percent on a four-hour one, the same halving pattern a half-life describes.

The mechanism is that a longer task is really a chain of more subtasks, and succeeding at the whole thing means succeeding at every link in it. Nothing about the model has to get worse for the success rate to fall; the same per-step reliability just gets multiplied against itself more times. Shortening the task, or checking in partway through instead of only at the end, changes how many multiplications you’re exposed to.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y