How do I catch a regression in an Agents SDK workflow?

Grade the traces the SDK already produces before building anything else. Every run has a span for each agent, model call, tool call, handoff, and guardrail, and OpenAI’s Graders can score those spans directly, which is enough to spot a workflow-level regression without a separate dataset.

Once you know what a passing trace looks like, promote the cases that matter into a fixed dataset and run it again with the Evals API after each change, comparing the new scores against the old ones case by case rather than as an average. Because a handoff and a tool call are separate, named spans, a comparison shows which piece of the workflow moved, not just that the final answer changed. LangGraph agents catch a graph-change regression the same way, rerunning a fixed dataset and diffing the two experiments node by node rather than eyeballing outputs.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y