How do you score turn-taking smoothness when a user interrupts the agent?
With a dedicated grader on the turn itself, because a transcript-only eval can score every one of a caller’s replies as accurate and still miss that the caller got talked over. IHBench, a 2026 benchmark on post-interruption recovery in voice agents, found GPT-family models correctly resume an utterance after a backchannel like “mm-hm” only 7% to 31% of the time, against 62% to 68% for Gemini 2.5. None of that shows up in the words: the resumed sentence still reads as a complete, plausible reply.
Grading it needs the trace to carry turn structure, who was speaking when, and where a backchannel or a real interruption landed inside an utterance. Instrumentation that records only the final transcript has nothing for a turn-taking grader to read, which is a gap in what got captured, not something a better grader can work around.
Two things get graded separately: how fast the agent yields the floor, and whether it resumes at the right point afterward. An agent can do one well and still fail the other.