Does SWE-bench Pro block an agent from looking up the fix online?
Yes. During the agent phase, the sandbox can reach the model endpoint and nothing else: no code hosts, no package indexes, no web-fetch tools. Scale AI verified that lock under the exact flags used for the real runs, with connections refused at connect time, and its V2 audit found no successful retrieval from a code host or module proxy during the public runs.
Even a network hole wouldn’t hand over much. Every task image is built from a sanitized bundle, so the fixing commit, stray refs, stashes, hooks, and test files simply aren’t in the repository the agent can see.
That closes off looking up the answer, not writing a fake one. SWE-bench Pro’s own V2 fixes a separate exploit, where a patch could tamper with the sandbox it was graded in rather than earn a passing test honestly; a benchmark’s scaffold decides what its score actually measures, and here the scaffold had to close two different doors, not one.