How do I replay a failed LangGraph run?
Call graph.get_state_history(config) on the run’s thread_id to get every checkpoint LangGraph saved for it, ordered most recent first, then pick the one where next names the node that failed. Invoking the graph again with that checkpoint’s id skips every node before it, since their results are already saved, and re-runs forward from exactly the state the graph had at that point.
That state doesn’t have to stay as it was. update_state() writes new values onto a chosen checkpoint and hands back a new checkpoint id, so you can change what the failing node received, an argument, a retrieved document, before running forward again. The original checkpoint isn’t touched, so the failing path stays there to compare against. This is real replay rather than a reconstruction, because the graph reruns against the same state object it actually had, not a fresh session built from a paraphrase of what happened, which is the same standard any agent’s failure replay has to meet, not just LangGraph’s.
sources
- LangGraph docs: Checkpointers fetched