How much has the top SWE-bench Pro score improved since launch?
The top public SWE-bench Pro score has nearly tripled since launch: from 23.3%, GPT-5’s score when Scale AI released the benchmark in September 2025, to 61.5% a year later, held by Meta’s Muse Spark 1.1 on Scale’s own public leaderboard.
That climb says more about how fast frontier models absorb a fixed, curated task set than about how well they’d do on code they’ve genuinely never seen. Scale built Pro from repositories no released model had trained on specifically to avoid that, and at launch the same models scored several points lower on the private commercial set than on the matched public one: 14.9% for GPT-5, 17.8% for Claude Opus 4.1. Production reliability is measured on your own traces, not a leaderboard climb, since the task set and scaffold behind this number both stayed fixed while only the models improved against them.