# Gaia

Gaia2 is Meta's test for personal-assistant agents, released in September 2025. The AI
operates a simulated phone with contacts, messages, email, calendar, and apps, and gets
tasks like "book the dinner and tell everyone who's coming." There are 1,120 scenarios.

The unusual part is that the clock keeps running. While the model is thinking, time passes
and things happen: a message arrives, a meeting moves, a friend replies. A verifier checks
every change the agent made against what should have happened, including whether it happened
in the right order and at the right time. It agrees with human graders 98% of the time. The
score is the share of scenarios passed on one attempt.

The scenarios are split by what they test: carrying out a sequence of actions, finding
information, spotting a task that's impossible or contradictory and asking instead of
guessing, adapting when the situation changes, working against a deadline, coping with tool
failures and irrelevant noise, and coordinating with another agent.

The best score at release was 42%, from GPT-5 at its highest reasoning setting, with the
best open model at 20%. The ambiguity and time splits are where scores fall hardest, and
they're the ones closest to how a real assistant fails: doing the wrong thing confidently,
or doing the right thing too late.

## Questions answered under this concept

- [Does a high Gaia2 score mean an agent will work in production?](https://tessary.ai/answers/gaia/does-a-high-gaia2-score-mean-an-agent-will-work-in-production)
- [How is Gaia2 scored?](https://tessary.ai/answers/gaia/how-is-gaia2-scored)
- [What does a Gaia2 score mean?](https://tessary.ai/answers/gaia/what-does-a-gaia2-score-mean)
- [What is Gaia2?](https://tessary.ai/answers/gaia/what-is-gaia2)

---

Source: https://tessary.ai/answers/gaia
All concepts: https://tessary.ai/answers
From Tessary, agent reliability for AI agents in production: https://tessary.ai
