# What does a SWE-bench Pro score mean?

A SWE-bench Pro score leans on requirement clarifications Scale writes into each task, not just raw repository comprehension: the paper's own ablation strips those clarifications out and top models collapse, GPT-5 (high) from 25.9% to 8.40%, Claude Opus 4.1 from 22.7% to 8.20%, on the identical bugs. Scale's stated reason is that without a human spelling out the exact requirement, the automated test verifier itself starts returning false negatives on a correct fix.

The score also caps what counts as an attempt: a single try, inside a fixed budget of 50 agent turns and $2 of inference cost per task. Production doesn't hand an agent a human's requirement clarification either, so [reliability there gets measured against your own live traces](/answers/agent-reliability/how-do-you-measure-agent-reliability-in-production), not a fixed task list with the hardest part of understanding the bug already solved for it.

---

Sources:
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? (arXiv:2509.16941): https://arxiv.org/abs/2509.16941 (fetched 2026-09-22)

Source: https://tessary.ai/answers/swe-bench/what-does-a-swe-bench-pro-score-mean
More on Swe bench: https://tessary.ai/answers/swe-bench
From Tessary, agent reliability for AI agents in production: https://tessary.ai
