Does a high SWE-bench Pro score mean an agent will work in production?
No. A 2026 audit called BenchJack found that SWE-bench Pro’s runner trusts a result file the agent’s own patch can rewrite: a submitted patch that drops an extra, untouched file into the repository can overwrite the evaluator’s own parser script before it ever reads the test output, so the patch fabricates a passing result for every test name the task expects, whether or not the underlying bug got fixed. Because the scoring rule only checks that the expected test names show up among the reported passes, invented names satisfy it as well as real ones, and the audit verified the exploit end to end across all 731 tasks in the version of the benchmark it tested.
SWE-bench Pro isn’t an outlier here: the same audit drove nine of the ten benchmarks it tested, including SWE-bench Verified through a different gap in the same trust boundary, to a near-perfect hack rate. A benchmark score belongs to the scaffold it ran in, not the model alone, and here the scaffold trusted evidence the agent’s own patch could write.