# Can a deterministic check prevent MAST failure modes without an LLM judge?

Mostly not yet, by the one attempt's own follow-up audit: a deterministic layer that checks agent handoffs against a written contract, no model involved, reported gains everywhere at first, then turned out to have credited itself with catches that were really its checker's own bugs, and once corrected the governed runs score below ungoverned ones in four of six workflows.

The layer, called Maat, validates each handoff against a versioned contract instead of grading output, a pass-or-fail check closer in spirit to [a deterministic pre-merge gate](/answers/deploy-gates/does-production-detection-replace-the-pre-deploy-gate) than to an LLM judge. Across 522 trials on six workflows with injected defects, the first run claimed gains of 2.9 to 26.5% everywhere. A hand review of all 94 halts it triggered found 35, 37%, were the validator misreading a valid handoff as broken.

Corrected for that, the signal is mixed: a halt that caught a real defect lifted scores 7.7 to 29.1 points in five workflows and cut cost 17 to 53%, but software development showed no gain, and counting every false alarm as lost work puts the governed arm behind doing nothing in most domains. It's a single-author preprint published days ago; treat the numbers as unsettled, not the idea itself.

---

Sources:
- "Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows" (arXiv:2609.34017): https://arxiv.org/abs/2609.34017 (fetched 2026-09-30)

Source: https://tessary.ai/answers/mast-taxonomy/can-a-deterministic-check-prevent-mast-failure-modes
More on Mast taxonomy: https://tessary.ai/answers/mast-taxonomy
From Tessary, agent reliability for AI agents in production: https://tessary.ai
