all answers

Frameworks and tooling

Langgraph evals

LangGraph is LangChain's library for building an agent as a graph. Each node is a function, edges decide which node runs next, and one state object is passed through the whole run. That structure is what makes a LangGraph agent easier to evaluate than a loop of model calls: every step has a name, and a failure can be pinned to a node rather than to "the agent".

Three things follow from it. LangGraph saves the state after every step, keyed by a thread id, and that is what makes pause, human approval, and resume work. It also means you can rewind a failed run to a checkpoint, change the state, and run it forward again with the exact context it had at the time, which is real replay rather than a reconstruction. Because each node is a unit, a check can be scoped to one node, so "the router picked the wrong branch" is a separate finding from "the answer was wrong". And a change to the graph is a diff, so when a regression appears after a change, the set of nodes that could have caused it is short.

Tracing is where teams get stuck. LangGraph does not emit OpenTelemetry spans on its own. It traces to LangSmith, and the way to send those same spans anywhere else is LangSmith's OTLP export: set LANGSMITH_TRACING and LANGSMITH_OTEL_ENABLED, and point OTEL_EXPORTER_OTLP_ENDPOINT at your collector.

9 questions

Answered, plainly.

How do I trace a LangGraph agent?LangGraph traces through LangSmith's OTLP export, not a built-in OTel exporter: set LANGSMITH_TRACING and LANGSMITH_OTEL_ENABLED, then point at your collector.answer →How do I grade one node of a LangGraph graph?Call the node directly through the compiled graph's own nodes dict and pass it to an evaluator the same way you'd pass the full graph.answer →How do I catch a regression caused by a graph change?Run the same eval dataset against the graph before and after the change and compare the two experiments; LangSmith keys both to the same cases automatically.answer →How do I replay a failed LangGraph run?Call get_state_history on the thread to list its checkpoints, find the one before the failing node, and re-invoke the graph from that checkpoint's id.answer →What's the difference between a LangGraph workflow and a LangGraph agent?A workflow's control flow is code you wrote in advance; an agent's control flow is a decision the model makes itself at runtime, path by path.answer →What's the difference between a LangGraph agent and an AWS Strands agent?A LangGraph agent still runs inside a graph you wrote; a default Strands agent has none, since AWS calls its single-agent loop 'model-driven,' though Strands also offers Graphs as an opt-in pattern.answer →What's the difference between a LangChain agent and a LangGraph agent?A LangChain agent built with create_agent already runs on LangGraph underneath; the real choice is that thin interface versus building the graph yourself.answer →How do I test a LangGraph agent that pauses for human approval?Resume the paused graph with Command(resume=...) on the same thread id and supply the approve, edit, reject, or respond value your test wants.answer →What's the difference between LangGraph and Deep Agents?LangGraph is the low-level graph runtime; Deep Agents is a batteries-included agent built on top of it, with planning, a virtual filesystem, and subagents wired in.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y