Are most agent errors caused by the agent's own logic?

No, most of what shows up as an error span isn’t the agent reasoning badly. Datadog’s State of AI Engineering 2026 report found that in February 2026, 5 percent of all LLM call spans reported an error, and 60 percent of those errors were exceeded rate limits, nothing in the agent’s own logic.

Framework and infrastructure code makes up a growing share of what runs between a prompt and a response, since agent framework adoption nearly doubled year over year in the same report. A rate limit, a provider timeout, or a dependency bump can produce a failure that looks identical to a reasoning error on a dashboard that only tracks pass or fail. Treating every error span as a model problem sends the investigation at the wrong layer. Check what actually threw before assuming the agent got the task wrong; a good share of the time, it never got the chance to.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y