# How do I catch a regression in an Agents SDK workflow?

Grade the traces the SDK already produces before building anything else. Every run has a span for each agent, model call, tool call, handoff, and guardrail, and OpenAI's Graders can score those spans directly, which is enough to spot a workflow-level regression without a separate dataset.

Once you know what a passing trace looks like, promote the cases that matter into a fixed dataset and run it again with the Evals API after each change, comparing the new scores against the old ones case by case rather than as an average. Because a handoff and a tool call are separate, named spans, a comparison shows which piece of the workflow moved, not just that the final answer changed. [LangGraph agents catch a graph-change regression the same way](/answers/langgraph-evals/how-do-i-catch-a-regression-caused-by-a-graph-change), rerunning a fixed dataset and diffing the two experiments node by node rather than eyeballing outputs.

---

Sources:
- OpenAI platform docs: Evaluate agent workflows: https://developers.openai.com/api/docs/guides/agent-evals (fetched 2026-09-19)

Source: https://tessary.ai/answers/openai-agents-sdk-evals/how-do-i-catch-a-regression-in-an-agents-sdk-workflow
More on Openai agents sdk evals: https://tessary.ai/answers/openai-agents-sdk-evals
From Tessary, agent reliability for AI agents in production: https://tessary.ai
