How do I catch a regression in an Agents SDK workflow?
Grade the traces the SDK already produces before building anything else. Every run has a span for each agent, model call, tool call, handoff, and guardrail, and OpenAI’s Graders can score those spans directly, which is enough to spot a workflow-level regression without a separate dataset.
Once you know what a passing trace looks like, promote the cases that matter into a fixed dataset and run it again with the Evals API after each change, comparing the new scores against the old ones case by case rather than as an average. Because a handoff and a tool call are separate, named spans, a comparison shows which piece of the workflow moved, not just that the final answer changed. LangGraph agents catch a graph-change regression the same way, rerunning a fixed dataset and diffing the two experiments node by node rather than eyeballing outputs.