Is an agent that correctly refuses a task a completion failure?
To a task-completion grader, yes, and that’s a real limit, not a bug to code around. The grader reads the trace for a state change, an order placed, a ticket resolved, a refund issued, and passes it when that state exists. A refusal the agent’s policy actually required and a refusal that’s the agent giving up mid-task produce the identical trace: no state change reached. The grader has nothing else to read, so it can’t tell which one happened.
Separating them needs a second check on the refusal’s own stated reason, not a better completion grader: did the policy actually prohibit this action, or did the agent just stall. Skip that second check on a policy-heavy workflow, refunds, account changes, anything an agent should sometimes decline, and every correct refusal drags the completion rate down next to every genuinely abandoned task.
It’s the mirror image of an agent that claims a task is done when it isn’t: one grader gets fooled by false confidence, the other by a true no.