Does an AI coding agent fix its own mistakes?

Rarely on its own. A study of 20,574 real coding-agent sessions found the agent almost never caught a mistake unprompted: a developer had to flag the problem 91.49% of the time before the agent went and fixed it, against just 2.99% where the agent self-corrected with no pushback at all, and 5.52% where the developer ended up fixing it directly instead. The misalignment took seven recurring forms: misreading the project, misreading intent, breaking a stated rule, overstepping its bounds, a bad implementation, a malformed command or tool call, and misreporting its own progress. Most of these episodes, 90.50%, cost extra effort and eroded trust rather than causing damage that couldn’t be undone.

That’s the rate for mistakes a person already caught. What happens to the ones nobody notices is a separate, harder question this study doesn’t answer, since it only counted episodes visible enough to draw a developer’s pushback in the first place.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y