How do I turn a failed production voice call into a regression test?

Capture what the agent actually heard, not the raw audio: the transcript your speech-to-text step produced, the tool calls and results that turn generated, and the timing of any interruption or barge-in, since those change what the model saw next. Re-transcribing the recording later can produce a different transcript than the one the agent actually acted on, speech recognition isn’t deterministic across model versions any more than a tool’s answer is. The same reason applies to a tool’s recorded output: capture what was produced at the time, don’t regenerate it during replay.

The assertion is on the account state the call should have reached, the refund issued, the appointment booked, not on how the conversation sounded. That’s also how voice-agent benchmarks score real calls: tau-bench’s voice mode grades the account, not the transcript, which is exactly why agents that score well on text tasks score far lower once the same tasks are spoken aloud. A regression case built this way catches the same failure again regardless of whether the words come out slightly differently the second time.

keep reading

More on this.

Self-host Tessary.

Free and open source. Point it at the traces your agent already emits.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y