How do I replay a failed LangGraph run?

Call graph.get_state_history(config) on the run’s thread_id to get every checkpoint LangGraph saved for it, ordered most recent first, then pick the one where next names the node that failed. Invoking the graph again with that checkpoint’s id skips every node before it, since their results are already saved, and re-runs forward from exactly the state the graph had at that point.

That state doesn’t have to stay as it was. update_state() writes new values onto a chosen checkpoint and hands back a new checkpoint id, so you can change what the failing node received, an argument, a retrieved document, before running forward again. The original checkpoint isn’t touched, so the failing path stays there to compare against. This is real replay rather than a reconstruction, because the graph reruns against the same state object it actually had, not a fresh session built from a paraphrase of what happened, which is the same standard any agent’s failure replay has to meet, not just LangGraph’s.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y