Can adding more tools make an agent worse?

Yes. EvoHarnessBench, a 2026 benchmark that grows a harness stage by stage while the model’s own weights stay fixed, found that expanding the harness alone can degrade performance on tasks the agent had already solved, a pattern the authors call harness-induced forgetting. Across 17 harness streams built from 802 tasks, deployment performance dropped 12.1% when tools were added, 13.8% when skills were added, and 46.4% when specialist agents were added, the largest of the three.

The old capability doesn’t disappear: the model can still do what it could before, but that behavior now has to be found and executed inside a bigger pool of tools, skills, or agents, which makes it harder to recover reliably. A harness grows for a reason, usually to cover a new case, but every addition raises the chance the agent picks the wrong tool for a task it already had covered.

Testing a harness change against the tasks it already passes, not just the one it was built for, catches this before it ships.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y