all answers

Agent reliability

Failure replay

Failure replay is the practice of reproducing a production failure with its full original context, so that a proposed fix can be verified against the case that actually happened.

It exists because agent failures are context-dependent. A failure comes from a specific conversation history, a specific set of tool results, and specific state. The same prompt with a paraphrased message or slightly different state will often behave differently, so a fix checked against an approximation of the failure can pass while the original case still fails.

Replaying a failure means capturing enough of the trace to re-run the failing turn faithfully: the message history the model saw, the tool outputs as they were returned at the time, and the configuration that was active. Tool outputs have to be recorded because the live systems behind them change state, and a later call would return different results. With that context preserved, a proposed fix runs against the real failing case, so its verdict reflects the actual failure.

The captured case also outlives the fix. Once a failure is reproducible, it can join an eval set as a regression case, so the same failure class is checked on every future change. This gives replay a second role beyond debugging: the record built to understand one failure becomes a standing check for its recurrence.

8 questions

Answered, plainly.

How do I replay a failed agent session turn by turn?Missed steps from an idle tab recover the same way any failed turn does: replay the exact message history and tool outputs the agent had at the time, in order.answer →What can replay not tell me?Replay tells you whether a given input and context reproduce an output, not why the model chose it, and not how often the failure happens across other sessions.answer →What does replay show that a trace viewer does not?A trace viewer shows what happened. Replay shows what would happen if you changed one thing, the prompt or a tool result, and re-ran the rest against the original context.answer →Why can't I just reproduce the failure locally?A local re-run rarely reproduces a production failure because the tool results and state that shaped it are gone by the time you retype the message and try again.answer →Why do tool outputs need to be captured instead of called again live?Tool outputs must be captured because the live systems behind them change state; calling a tool again during replay can return a different answer than the agent originally saw.answer →Why does a fixed agent bug come back weeks later?A fixed bug comes back when the fix is verified against a rewritten description of the failure instead of the replayed case, so nothing keeps checking for that exact case.answer →How do I turn a failed production tool call into a regression test?Capture the failing turn with its recorded tool outputs and the config in force, pin the end state it should have reached, and run it as a deterministic check on every change.answer →How do I turn a failed production voice call into a regression test?Capture the transcript, tool calls, and interruption timing the agent actually acted on, then assert on the account state the call should have reached.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y