# What changed in the latest version of Terminal-Bench?

Terminal-Bench 4.0, released August 2026, dropped eight of 3.0's 74 tasks, two each for being saturated, refusal-prone, having a publicly available solution, or carrying an unresolved quality or platform issue, and fixed 19 more by rewriting their instructions, environments, or verifiers.

None of that changes what a score means, only what it's measuring against. A task every frontier model already solves five times out of five stops separating models from each other, so keeping it in the set just adds noise around a fixed point everyone hits. The 19 fixes work the other way: they close gaps where a task failed an agent for a reason that had nothing to do with the agent's actual work, like a broken verifier or an ambiguous instruction. [A benchmark score is reported for the model-and-harness pair, not the model alone](/answers/agent-benchmarks/what-is-an-agent-benchmark), so a rewritten verifier can shift a given pair's score even when the agent's real behavior on that task hasn't changed at all.

---

Sources:
- Terminal-Bench 4.0 (tbench.ai): https://www.tbench.ai/news/terminal-bench-4-0 (fetched 2026-09-17)

Source: https://tessary.ai/answers/terminal-bench/what-changed-in-the-latest-version-of-terminal-bench
More on Terminal bench: https://tessary.ai/answers/terminal-bench
From Tessary, agent reliability for AI agents in production: https://tessary.ai
