What does a Gaia2 score mean?

A Gaia2 score is an average across seven very different task splits, and that average hides how uneven agents actually are. GPT-5 at its highest reasoning setting posts 79.6% on search tasks and 69.2% on execution tasks, the two easiest splits, then drops to 0.0% on Time, the split where a scenario carries a real deadline. Its headline 42.1% overall is mostly the easy splits carrying the hard ones.

The Time split’s zero isn’t a policy failure. Meta’s own ablation reruns the same tasks in an instant-time mode that removes real inference latency, and GPT-5 (high) jumps from 0.0% to 34.4%: a reasoning model can spend more wall-clock time thinking than the scenario’s deadline allows, so a plan that would have worked still counts as a miss. Noise, where a scenario adds irrelevant distractions, is the other weak split, with most models scoring under 20% there.

Read one Gaia2 percentage as a blend of splits an agent may be nowhere near equally good at, not a single policy score. Tessary’s duration_drift classifier watches the same kind of latency sensitivity in production, against each call site’s own history rather than a scenario’s fixed deadline.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y