all answers

Frameworks and tooling

Mcp tool reliability

The Model Context Protocol (MCP) is a standard way to give a model tools that live in a separate process or service. The agent connects to an MCP server, asks what tools it offers, and calls them. What this changes for reliability is ownership: the tools your agent depends on are now someone else's deployment, updated on someone else's schedule.

Four things go wrong because of that.

Errors come back in two shapes, and servers mix them up. A tool that ran and failed is supposed to return a normal result marked as an error, with the reason in the content, so the model can read it and try something else. A protocol problem, like an unknown tool or a malformed request, comes back as a protocol error the model never sees. When a server reports a tool failure as a protocol error, the agent has no way to recover from it.

Tool definitions change under you. Clients cache the list of tools, a server is supposed to announce when its tools change, and a client that misses that announcement builds arguments for a tool that no longer looks like that.

Slow servers become agents that never return. Every request needs a timeout, and the timeout has to hold even while the server keeps sending progress updates.

Cancellation differs by transport. Over HTTP, closing the stream is the cancel. Over stdio, it's an explicit message, and both sides have to cope with a response that arrives after the cancel anyway.

Evaluating an MCP server means testing each of these on purpose: return a failed tool result, change a schema mid-session, stall a response, cancel a call, and check what the agent does.

10 questions

Answered, plainly.

Should an MCP tool failure be a protocol error or a failed result?A tool that ran and failed should return isError: true so the model sees the reason; a protocol error is only for a request the server never attempted.answer →What happens when an MCP server changes its tool schema?A "stale catalog" error means the client cached tool schemas from an old tools/list call; the tool keeps getting called against the old version until it refreshes.answer →How do I set a timeout on an MCP tool call?Set a per-request timeout in your client SDK and cancel on expiry; the spec lets a progress notification reset the clock, but a hard maximum should still apply.answer →How do I cancel an in-flight MCP tool call?Send a notifications/cancelled message naming the request id; the server stops and skips a response, but a reply can still arrive if it crosses the cancel.answer →How do I evaluate an MCP server's reliability?Test the four ways an MCP server fails on purpose: a failed tool result, a schema that changes mid-session, a stalled response, and a cancel, each checked.answer →What's the difference between MCP and A2A?MCP connects an agent to tools that run in someone else's deployment; A2A connects one agent to another so they can delegate a task, and both usually run together.answer →What's the difference between MCP and Claude Skills?MCP gives Claude access to external tools and data; Skills teach it how to use what it has, loading procedural instructions only when a task calls for them.answer →What's the difference between MCP and a plain API call?Both fetch data or trigger an action. A direct API call uses a contract written into your code; MCP gets a tool schema from a server at runtime.answer →What's the difference between MCP and RAG?RAG retrieves material for an answer; MCP is how an agent discovers and calls a tool. A retriever can use MCP, but neither term decides who runs it.answer →What's the difference between MCP and LangChain?MCP is an open protocol for exposing tools to any compatible client; LangChain is an orchestration library that, since v1.4, ships a built-in client for consuming MCP servers.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y