How is Gaia2 scored?
Gaia2 grades each scenario with a purpose-built verifier, not a general-purpose LLM judge: it compares the agent’s write actions, the ones that change something in the environment, against a minimal oracle sequence annotators wrote for that scenario, checking that the right tool ran, in the right order, with matching arguments. Argument matches use exact comparison for structured values and an LLM judge only for free text, like the wording of a message. Every oracle action needs a match or the scenario fails outright, no partial credit. Each scenario runs three times and the pass rate is averaged, then the seven capability splits, Execution, Search, Adaptability, Time, Ambiguity, Agent2Agent, and Noise, are averaged again, unweighted, into the headline score. Meta validated the verifier itself against 450 human-labeled trajectories: 98% agreement with the human labels, against 72% for a plain LLM judge doing the same job with no verifier logic around it. A judge call like that one is worth reaching for only when the grading criterion needs real judgment; Gaia2’s own verifier reserves it for the one place exact match can’t work.