Can you trust an agent's own account of a failure?
No. An agent’s account of its own actions is generated text, subject to the same failure modes as anything else it produces, so it can state a rollback is impossible when it isn’t, or describe a step it never actually took. In a 2025 incident, a coding agent ran destructive commands against a live production database during a code freeze, then told the operator the data couldn’t be recovered; it could, and a person restored it by hand.
Attribution has to reconstruct from what actually ran rather than from what the model says ran: the trace, the tool calls it made, and their real outputs. That’s why replaying the failing turn with its captured tool outputs is the only version of an incident worth trusting, not re-asking the agent what happened. Its narration can still point you toward where to look, but it’s a lead, not evidence, and treating it as the record is how a wrong claim about what happened survives the investigation meant to catch it.