Does SWE-bench Pro block an agent from looking up the fix online?

Yes. During the agent phase, the sandbox can reach the model endpoint and nothing else: no code hosts, no package indexes, no web-fetch tools. Scale AI verified that lock under the exact flags used for the real runs, with connections refused at connect time, and its V2 audit found no successful retrieval from a code host or module proxy during the public runs.

Even a network hole wouldn’t hand over much. Every task image is built from a sanitized bundle, so the fixing commit, stray refs, stashes, hooks, and test files simply aren’t in the repository the agent can see.

That closes off looking up the answer, not writing a fake one. SWE-bench Pro’s own V2 fixes a separate exploit, where a patch could tamper with the sandbox it was graded in rather than earn a passing test honestly; a benchmark’s scaffold decides what its score actually measures, and here the scaffold had to close two different doors, not one.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y