How do I evaluate an MCP server's reliability?
Test the four ways an MCP server can fail on purpose, rather than trusting its advertised uptime: return a failed tool call, change a tool’s schema mid-session, stall a response, and cancel a call, then check what your agent does in each case.
Each check targets one behavior your agent depends on. Send back a result marked isError: true and confirm the model reads the reason instead of treating the call as a success. Change a tool’s parameters between calls without the notifications/tools/list_changed message and check whether the agent still builds arguments for the old shape. Hold a response open past your timeout and confirm the agent gives up instead of hanging. Send a cancellation mid-call and confirm the server stops instead of finishing and replying anyway.
The MCP Inspector’s CLI mode scripts the failed-tool-result check: call tools/call and read the exit code, 5 on isError: true. The other two need a session the CLI doesn’t hold open, since each run connects, sends one request, and exits; test those through the web client or your agent itself. How reliable MCP tool calls are in practice is the number this testing is trying to catch before a stress test finds it.