How do I replay a failed agent session turn by turn?

When an agent’s tab goes idle and misses a run of steps, you recover them the same way you replay any failed turn: feed the model the exact message history and tool outputs it had at that point, in order, instead of letting live tools answer again. Capture three things per turn: the messages up to that point, the tool call and its recorded output, and the configuration, model, prompt version, settings, that was active. Replay the turns in order, holding tool outputs fixed, so each step sees exactly what the original run saw. If a step calls a tool again instead of reading the recorded result, you’re not replaying, you’re running a new session that happens to start the same way and can diverge from turn one. Doing it turn by turn, rather than replaying the whole session as one unit, lets you stop at the turn where things went wrong and inspect the model’s state right there.

See also: What is regression testing for an AI agent?

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y