# What does a Gaia2 score mean?

A Gaia2 score is an average across seven very different task splits, and that average hides how uneven agents actually are. GPT-5 at its highest reasoning setting posts 79.6% on search tasks and 69.2% on execution tasks, the two easiest splits, then drops to 0.0% on Time, the split where a scenario carries a real deadline. Its headline 42.1% overall is mostly the easy splits carrying the hard ones.

The Time split's zero isn't a policy failure. Meta's own ablation reruns the same tasks in an instant-time mode that removes real inference latency, and GPT-5 (high) jumps from 0.0% to 34.4%: a reasoning model can spend more wall-clock time thinking than the scenario's deadline allows, so a plan that would have worked still counts as a miss. Noise, where a scenario adds irrelevant distractions, is the other weak split, with most models scoring under 20% there.

Read one Gaia2 percentage as a blend of splits an agent may be nowhere near equally good at, not a single policy score. [Tessary's duration_drift classifier](/answers/tessary-duration-drift/what-is-tessarys-duration-drift-classifier) watches the same kind of latency sensitivity in production, against each call site's own history rather than a scenario's fixed deadline.

---

Sources:
- ARE: Scaling Up Agent Environments and Evaluations (arXiv:2509.17158): https://arxiv.org/abs/2509.17158 (fetched 2026-09-18)

Source: https://tessary.ai/answers/gaia/what-does-a-gaia2-score-mean
More on Gaia: https://tessary.ai/answers/gaia
From Tessary, agent reliability for AI agents in production: https://tessary.ai
