# Does a better harness matter more than a better model?

Sometimes: tuning the harness has moved a fixed model's score as much as switching to a better model did, but the effect varies so much by model that there's no general rule. On a hard subset of tau2-bench, LangChain's [per-model harness profiles](https://www.langchain.com/blog/tuning-deep-agents-different-models) lifted GPT 5.3 Codex from 33% to 53% and Claude Opus 4.7 from 43% to 53%, so tuning closed the whole gap between the two models. LangChain also [took its coding agent](https://www.langchain.com/blog/improving-deep-agents-with-harness-engineering) from 52.8% to 66.5% on Terminal-Bench 2.0 with the model fixed at gpt-5.2-codex, changing mainly prompts, middleware, and per-stage reasoning budget.

Other harness swaps move far less. On the verified [Terminal-Bench 2.1 leaderboard](https://www.tbench.ai/leaderboard/terminal-bench/2.1), Claude Opus 4.7 at max effort scores 68.9% inside Claude Code and 66.1% inside the Terminus 2 reference agent, while Gemini 3.1 Pro at high effort moves 0.2 points between Gemini CLI and Terminus 2. Part of the spread is that a harness gets tuned to one model's habits, which is also why [a leaderboard score doesn't carry over to your setup](/answers/agent-benchmarks/what-is-an-agent-benchmark).

All of this comes from coding and tool-use benchmarks, much of it reported by teams tuning their own harness, so none of it measures your agent on your traffic.

---

Sources:
- LangChain, "Improving Deep Agents with harness engineering" (February 17, 2026): https://www.langchain.com/blog/improving-deep-agents-with-harness-engineering (fetched 2026-10-03)
- LangChain, "Tuning Deep Agents to Work Well with Different Models" (April 29, 2026): https://www.langchain.com/blog/tuning-deep-agents-different-models (fetched 2026-10-03)
- Terminal-Bench 2.1 leaderboard (tbench.ai): https://www.tbench.ai/leaderboard/terminal-bench/2.1 (fetched 2026-10-03)

Source: https://tessary.ai/answers/agent-harness/does-a-better-harness-matter-more-than-a-better-model
More on Agent harness: https://tessary.ai/answers/agent-harness
From Tessary, agent reliability for AI agents in production: https://tessary.ai
