How is Gaia2 scored?

Gaia2 grades each scenario with a purpose-built verifier, not a general-purpose LLM judge: it compares the agent’s write actions, the ones that change something in the environment, against a minimal oracle sequence annotators wrote for that scenario, checking that the right tool ran, in the right order, with matching arguments. Argument matches use exact comparison for structured values and an LLM judge only for free text, like the wording of a message. Every oracle action needs a match or the scenario fails outright, no partial credit. Each scenario runs three times and the pass rate is averaged, then the seven capability splits, Execution, Search, Adaptability, Time, Ambiguity, Agent2Agent, and Noise, are averaged again, unweighted, into the headline score. Meta validated the verifier itself against 450 human-labeled trajectories: 98% agreement with the human labels, against 72% for a plain LLM judge doing the same job with no verifier logic around it. A judge call like that one is worth reaching for only when the grading criterion needs real judgment; Gaia2’s own verifier reserves it for the one place exact match can’t work.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y