Is my agent getting worse because of the model or the harness?

It can be either, and you tell them apart by swapping one at a time: rerun the inputs that got worse, several times each, on the previous model with today’s harness, then on today’s model with the previous harness, and whichever swap brings the old results back points at the cause.

In April 2026, Anthropic traced weeks of complaints that Claude had gotten worse to three changes in its own harness, including a default reasoning effort lowered from high to medium and a bug that dropped Claude’s earlier reasoning on every turn, and wrote that “The API was not impacted.” The changes reached Claude Code, the Claude Agent SDK, and Claude Cowork, so a team building on a provider’s harness can be hit by a change it never made.

Within the harness, finding the specific change still means testing each candidate against the failing traces. A replay only catches a harness bug if it recreates the state the bug reacts to: Anthropic’s reasoning bug started once a session had sat idle for over an hour, and its own evals didn’t reproduce the issues at first.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y