Why does an AI agent's success rate drop so much on a task that takes longer?
Because each extra step or minute of work carries roughly the same chance of failure, and those chances compound the way radioactive decay does. A 2025 analysis of agent benchmark data found that treating each minute a task would take a human as carrying a constant hazard of failure predicts real success rates surprisingly well: an agent that completes a one-hour task with 50 percent probability drops to about 25 percent on a two-hour task and about 6 percent on a four-hour one, the same halving pattern a half-life describes.
The mechanism is that a longer task is really a chain of more subtasks, and succeeding at the whole thing means succeeding at every link in it. Nothing about the model has to get worse for the success rate to fall; the same per-step reliability just gets multiplied against itself more times. Shortening the task, or checking in partway through instead of only at the end, changes how many multiplications you’re exposed to.