# What changed in the latest version of CursorBench?

CursorBench 4.0, released September 10, 2026, added long-horizon tasks: multi-step edits, refactors, investigations, working out what an underspecified request actually meant, managing a job across more than one session, and sticking to an existing design, on top of the shorter edit-refactor-bugfix tasks earlier versions scored. Each prior version widened the set the same way: 3.1 added codebase understanding and code review, 3.2 added instruction-following and advanced tool use, and 4.0 is the first built around tasks that span more than one sitting.

Cursor's own leaderboard page says 4.0 scores aren't comparable to the retired 3.2 set, since the task categories themselves changed rather than the same categories just getting harder, [the more common way a benchmark version bump lowers scores](/answers/agent-benchmarks/why-do-benchmark-scores-drop-when-a-new-version-is-released). The top model on 4.0 at publication is Claude Fable 5.1 (Max) at 51.8%, and Cursor still hasn't published how correctness is graded or how many tasks the set contains.

---

Sources:
- CursorBench (cursor.com): https://cursor.com/cursorbench (fetched 2026-09-18)

Source: https://tessary.ai/answers/cursorbench/what-changed-in-the-latest-version-of-cursorbench
More on Cursorbench: https://tessary.ai/answers/cursorbench
From Tessary, agent reliability for AI agents in production: https://tessary.ai
