How do I grade one node of a LangGraph graph?

Call the node directly instead of running the whole graph. A compiled LangGraph exposes each node on its own nodes dict, so you chain a small function that builds that node’s input state onto app.nodes["node_name"] and pass the result to an evaluator the same way you’d pass the full graph; LangSmith’s aevaluate() takes either target interchangeably. The node runs alone, so grading it costs only that node’s own model or tool call, not the rest of the graph’s.

Do this separately from grading the final answer, because a node can do its own job correctly and still leave the graph wrong: a downstream node receives a state the graph never should have handed it, even though the node itself ran fine on the input it got. That’s the same handoff failure any multi-agent system risks, and node-level checks alone won’t catch it, since each node is graded on the input it was given, not on whether that was the right input to give it. A full suite needs both: a node check for whether the step did its job, and a graph-level check for whether the right state reached it.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y