What does a SWE-bench Pro score mean?

A SWE-bench Pro score leans on requirement clarifications Scale writes into each task, not just raw repository comprehension: the paper’s own ablation strips those clarifications out and top models collapse, GPT-5 (high) from 25.9% to 8.40%, Claude Opus 4.1 from 22.7% to 8.20%, on the identical bugs. Scale’s stated reason is that without a human spelling out the exact requirement, the automated test verifier itself starts returning false negatives on a correct fix.

The score also caps what counts as an attempt: a single try, inside a fixed budget of 50 agent turns and $2 of inference cost per task. Production doesn’t hand an agent a human’s requirement clarification either, so reliability there gets measured against your own live traces, not a fixed task list with the hardest part of understanding the bug already solved for it.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y