Is an agent that correctly refuses a task a completion failure?

To a task-completion grader, yes, and that’s a real limit, not a bug to code around. The grader reads the trace for a state change, an order placed, a ticket resolved, a refund issued, and passes it when that state exists. A refusal the agent’s policy actually required and a refusal that’s the agent giving up mid-task produce the identical trace: no state change reached. The grader has nothing else to read, so it can’t tell which one happened.

Separating them needs a second check on the refusal’s own stated reason, not a better completion grader: did the policy actually prohibit this action, or did the agent just stall. Skip that second check on a policy-heavy workflow, refunds, account changes, anything an agent should sometimes decline, and every correct refusal drags the completion rate down next to every genuinely abandoned task.

It’s the mirror image of an agent that claims a task is done when it isn’t: one grader gets fooled by false confidence, the other by a true no.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y