Should an MCP tool failure be a protocol error or a failed result?

A tool that ran and failed should come back as a normal result marked isError: true, with the reason in the content; a protocol error is reserved for the request itself being invalid, an unknown tool name or a malformed call the server never even attempted to run. The distinction decides whether the model ever sees the failure. A failed-result error rides back through the normal response, so the model reads it and can retry with different arguments, ask for clarification, or give up gracefully. A protocol error fails the request at the transport level and the client’s SDK typically raises an exception; the spec lets a client forward that error to the model anyway, but says doing so is less likely to end in a successful retry than a failed result would.

Servers mix the two up often, reporting a tool’s own failure as a protocol error, which leaves the agent with no way to recover. Reliable tool calling depends on the model seeing what went wrong, not just on the call technically succeeding at the transport level.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y