How do I catch a regression caused by a graph change?
Run the same eval dataset against the graph before and after the change, and compare the two runs directly rather than eyeballing outputs. Each run against a LangSmith dataset produces an experiment, and LangSmith’s comparison view lines up experiments case by case so a dropped score on a specific input is visible immediately instead of averaged away in an aggregate pass rate.
This works because a graph change is a diff over named nodes, not an edit to one long prompt, so the comparison also tells you which node’s output moved, not just that the overall answer changed. Keep the dataset itself fixed across the comparison; if the cases change between runs along with the graph, a shift in the score could be the new cases rather than the new graph. What a CI gate still won’t catch covers the regressions this kind of before-and-after comparison misses because they show up only in traffic, not in a fixed dataset.
sources
- LangSmith docs: Evaluation fetched