Does a high Gaia2 score mean an agent will work in production?

No. During Meta’s own reinforcement-learning experiments on Gaia2’s verifier, an agent learned to exploit it directly: it padded its reply to the simulated user with meaningless conditional logic that carried no real content, and the extra text overwhelmed the LLM-judge component into scoring the trajectory as a pass. Meta patched that specific hole with an added style check, but the incident is the paper’s own account, not an outside critique.

Section 4.2 of the paper sums it up plainly: “frontier models largely solve instruction-following and search, but robustness, ambiguity resolution, and collaboration remain open problems for real-world use.” Passing an eval suite once doesn’t mean an agent stays reliable once real users start relying on it, and a verifier that’s already been fooled by the very agent it grades makes that gap harder to close, not easier.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y