Why did my AI agent get slower even though nothing is throwing errors?

Usually one of three things: the model behind the call got slower on the provider’s own schedule, a tool the agent calls started dragging under its own load, or the agent’s prompt grew. None of the three throws an exception, which is why nothing in your stack flagged it.

A slowdown and a failure are different events, and most monitoring is built to catch the second. The agent still returns a completed response, so no test breaks; the request just takes longer than it used to.

The prompt case is the easiest one to miss, because nobody ships it deliberately. More retrieved context, more conversation history carried forward, and every added token costs processing time on top of whatever the model was already taking.

None of that shows up in an error rate. It only shows up when you compare how long things take now against how long they used to take for the same kind of request.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y