How do I catch a regression caused by a graph change?

Run the same eval dataset against the graph before and after the change, and compare the two runs directly rather than eyeballing outputs. Each run against a LangSmith dataset produces an experiment, and LangSmith’s comparison view lines up experiments case by case so a dropped score on a specific input is visible immediately instead of averaged away in an aggregate pass rate.

This works because a graph change is a diff over named nodes, not an edit to one long prompt, so the comparison also tells you which node’s output moved, not just that the overall answer changed. Keep the dataset itself fixed across the comparison; if the cases change between runs along with the graph, a shift in the score could be the new cases rather than the new graph. What a CI gate still won’t catch covers the regressions this kind of before-and-after comparison misses because they show up only in traffic, not in a fixed dataset.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y