Why does the same model score differently in different agent harnesses?
Because the harness decides what the model actually sees at each step, and that alone can move a score with the weights never touched. A 2026 study compared two configurations of the same harness: the control kept the full conversation in time order, and the treatment bundled two changes together, mechanically shortening older tool results as context filled and responding to repeated or stalled work. On a 169-task SWE-bench Verified slice run under a tight 20,480-token window, the treatment raised the model’s mean per-task fail-to-pass fraction, a continuous partial-credit score, from 28% to 49%, and raised complete solutions, a separate binary count, from 43 to 72.
The study is a single-author preprint and measures the combined effect of that treatment package, not any one rule inside it; the paper itself leaves isolating the detector response’s own contribution to future work. Applying the same frozen package to three other models, with no retuning, raised their scores too, so the effect was not specific to one model’s quirks.
A benchmark score is reported for the model-and-harness pair run together, and this is why: change only the harness, and the pair produces a different number.