What is regression testing for an AI agent?
Regression testing for an agent is re-running a fixed set of cases, each with an expected behavior, after every change, to catch behavior that used to work and stopped.
It differs from a unit test in what it asserts. The same input doesn’t produce the same output twice, so a case is judged rather than asserted, and the result is a pass rate over repeated runs instead of one green tick. A passing unit test doesn’t rule a regression out: it checks the code path, not what the model did once it got there.
The cases come from production. A failure that already happened is a real failure rather than a guess at one, and turning a failed turn into a standing check is how the set grows.
What it cannot catch is a change it has nothing to run against, since a regression can arrive with an empty diff.