Does a correct multi-agent answer mean nothing broke?
No, and which rule breaks decides how often it slips through. A study replayed multi-agent runs with one coordination rule broken at a time: whether information reaches the agent that needs it (routing), whether accepted evidence is weighted correctly against duplicate reports (admission), whether shared state stays current, and whether a correct decision gets carried out (action). Breaking routing, state, or action always produced a visibly wrong answer, in all 72 trials each. Breaking admission, treating two copies of one report as independent confirmation instead of catching the duplicate, still let the system land on the correct answer in 43 of 72 trials, with nothing in the final answer or the transcript showing anything had gone wrong.
Diagnosing the break afterward needed the same access the system itself did: reading only the public output caught 26% of the broken rules; reading each framework’s own internal records, the admission log, the state before and after an update, raised that to 63%. Grading each agent’s output separately misses the same kind of break, since a run that lands on the right answer gives no sign its coordination already failed. It’s a single-author preprint three days old, early evidence rather than settled.