Does a better harness matter more than a better model?
Sometimes: tuning the harness has moved a fixed model’s score as much as switching to a better model did, but the effect varies so much by model that there’s no general rule. On a hard subset of tau2-bench, LangChain’s per-model harness profiles lifted GPT 5.3 Codex from 33% to 53% and Claude Opus 4.7 from 43% to 53%, so tuning closed the whole gap between the two models. LangChain also took its coding agent from 52.8% to 66.5% on Terminal-Bench 2.0 with the model fixed at gpt-5.2-codex, changing mainly prompts, middleware, and per-stage reasoning budget.
Other harness swaps move far less. On the verified Terminal-Bench 2.1 leaderboard, Claude Opus 4.7 at max effort scores 68.9% inside Claude Code and 66.1% inside the Terminus 2 reference agent, while Gemini 3.1 Pro at high effort moves 0.2 points between Gemini CLI and Terminus 2. Part of the spread is that a harness gets tuned to one model’s habits, which is also why a leaderboard score doesn’t carry over to your setup.
All of this comes from coding and tool-use benchmarks, much of it reported by teams tuning their own harness, so none of it measures your agent on your traffic.