# How is SWE-bench Pro scored?

SWE-bench Pro scores a task as resolved only if the agent's patch makes the repository's held-out test suite pass: the fail2pass tests that confirm the reported bug is actually fixed, and the pass2pass tests that confirm nothing else broke. There's no partial credit and no second attempt: each task gets one try, capped at 50 agent turns and $2 of inference cost, and the reported score, Pass@1, is just the share of tasks resolved within whichever subset you're scoring, 731 in the public release, 276 in the commercial one. The test suite behind that pass/fail call goes through its own verification before a task ever ships: an automatic check that the environment builds and the tests aren't flaky, then a human reviewer checks every fail2pass test for relevance to the task and drops any that are too broad, dropping the task itself if none of its tests survive that bar. That review keeps the grading strict, not portable: [a benchmark score doesn't transfer to your own agent](/answers/agent-benchmarks/what-is-an-agent-benchmark), because each fail2pass test still expects the same interface and behavior its reviewer signed off on, whatever else a different correct fix might look like.

---

Sources:
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? (arXiv:2509.16941): https://arxiv.org/abs/2509.16941 (fetched 2026-09-27)

Source: https://tessary.ai/answers/swe-bench/how-is-swe-bench-pro-scored
More on Swe bench: https://tessary.ai/answers/swe-bench
From Tessary, agent reliability for AI agents in production: https://tessary.ai
