What does a SWE-bench Pro score mean?
A SWE-bench Pro score leans on requirement clarifications Scale writes into each task, not just raw repository comprehension: the paper’s own ablation strips those clarifications out and top models collapse, GPT-5 (high) from 25.9% to 8.40%, Claude Opus 4.1 from 22.7% to 8.20%, on the identical bugs. Scale’s stated reason is that without a human spelling out the exact requirement, the automated test verifier itself starts returning false negatives on a correct fix.
The score also caps what counts as an attempt: a single try, inside a fixed budget of 50 agent turns and $2 of inference cost per task. Production doesn’t hand an agent a human’s requirement clarification either, so reliability there gets measured against your own live traces, not a fixed task list with the hardest part of understanding the bug already solved for it.