Is my agent getting worse because of the model or the harness?
It can be either, and you tell them apart by swapping one at a time: rerun the inputs that got worse, several times each, on the previous model with today’s harness, then on today’s model with the previous harness, and whichever swap brings the old results back points at the cause.
In April 2026, Anthropic traced weeks of complaints that Claude had gotten worse to three changes in its own harness, including a default reasoning effort lowered from high to medium and a bug that dropped Claude’s earlier reasoning on every turn, and wrote that “The API was not impacted.” The changes reached Claude Code, the Claude Agent SDK, and Claude Cowork, so a team building on a provider’s harness can be hit by a change it never made.
Within the harness, finding the specific change still means testing each candidate against the failing traces. A replay only catches a harness bug if it recreates the state the bug reacts to: Anthropic’s reasoning bug started once a session had sat idle for over an hour, and its own evals didn’t reproduce the issues at first.