What does a Gaia2 score mean?
A Gaia2 score is an average across seven very different task splits, and that average hides how uneven agents actually are. GPT-5 at its highest reasoning setting posts 79.6% on search tasks and 69.2% on execution tasks, the two easiest splits, then drops to 0.0% on Time, the split where a scenario carries a real deadline. Its headline 42.1% overall is mostly the easy splits carrying the hard ones.
The Time split’s zero isn’t a policy failure. Meta’s own ablation reruns the same tasks in an instant-time mode that removes real inference latency, and GPT-5 (high) jumps from 0.0% to 34.4%: a reasoning model can spend more wall-clock time thinking than the scenario’s deadline allows, so a plan that would have worked still counts as a miss. Noise, where a scenario adds irrelevant distractions, is the other weak split, with most models scoring under 20% there.
Read one Gaia2 percentage as a blend of splits an agent may be nowhere near equally good at, not a single policy score. Tessary’s duration_drift classifier watches the same kind of latency sensitivity in production, against each call site’s own history rather than a scenario’s fixed deadline.