Can adding more tools make an agent worse?
Yes. EvoHarnessBench, a 2026 benchmark that grows a harness stage by stage while the model’s own weights stay fixed, found that expanding the harness alone can degrade performance on tasks the agent had already solved, a pattern the authors call harness-induced forgetting. Across 17 harness streams built from 802 tasks, deployment performance dropped 12.1% when tools were added, 13.8% when skills were added, and 46.4% when specialist agents were added, the largest of the three.
The old capability doesn’t disappear: the model can still do what it could before, but that behavior now has to be found and executed inside a bigger pool of tools, skills, or agents, which makes it harder to recover reliably. A harness grows for a reason, usually to cover a new case, but every addition raises the chance the agent picks the wrong tool for a task it already had covered.
Testing a harness change against the tasks it already passes, not just the one it was built for, catches this before it ships.