How do I turn a failed production tool call into a regression test?
Capture the failing turn, its message history, the tool outputs as they were recorded, and the config in force, pin the end state it should have reached, and run it as a deterministic check on every change.
Replay the captured tool output rather than calling the tool again, because the systems behind a tool keep moving. A fresh call turns the case into a new session that happens to start the same way.
The assertion is the exact condition that broke: a missing field, a wrong status, a step the agent claimed and never took. That is a comparison, not a judgment, so no LLM judge is involved and the check either passes or it does not.
trace_id 01JB6PZ4... turn 4
tool billing.get_invoice
captured {"status": "void", "amount_due": 0}
assert reply.states_amount_due == 0
The case tests the failure you saw, not the ones you did not. It is a standing check against that bug coming back, and one of two sources an eval dataset draws on.