How is SWE-bench Pro scored?

SWE-bench Pro scores a task as resolved only if the agent’s patch makes the repository’s held-out test suite pass: the fail2pass tests that confirm the reported bug is actually fixed, and the pass2pass tests that confirm nothing else broke. There’s no partial credit and no second attempt: each task gets one try, capped at 50 agent turns and $2 of inference cost, and the reported score, Pass@1, is just the share of tasks resolved within whichever subset you’re scoring, 731 in the public release, 276 in the commercial one. The test suite behind that pass/fail call goes through its own verification before a task ever ships: an automatic check that the environment builds and the tests aren’t flaky, then a human reviewer checks every fail2pass test for relevance to the task and drops any that are too broad, dropping the task itself if none of its tests survive that bar. That review keeps the grading strict, not portable: a benchmark score doesn’t transfer to your own agent, because each fail2pass test still expects the same interface and behavior its reviewer signed off on, whatever else a different correct fix might look like.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y