How do multi-agent systems actually fail?

Three root causes, and the largest is a bad instruction, not a bad model: the system was told the wrong thing, context broke down between agents, or nobody checked the output. The MAST taxonomy, built from 1,600+ annotated traces across seven multi-agent frameworks, sorts 14 failure modes into specification issues (41.77%), inter-agent misalignment (36.94%), and task verification failures (21.30%).

None of the three is a model-accuracy problem in the sense of the model getting a fact wrong. They’re coordination and process failures: the wrong instruction going out, the wrong information crossing a seam, or no one checking the output at all. That’s also why per-agent grading misses most of it: a grader that checks each agent’s own output in isolation has nothing to say about a specification that was ambiguous from the start or a handoff nobody verified.

What is the MAST taxonomy? breaks down the full classification these three root causes come from.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y