What's the difference between tool calling and code execution?
Ordinary tool calling puts every call and its full result through the model’s own context, one round trip per tool: the model emits one call, waits, reads the result, and decides what to do next; code execution instead has the model write a short program, in a sandboxed runtime, that imports the available tools as functions and calls several of them directly, so only the code’s own output returns to the model’s context, not each tool’s raw response along the way.
Anthropic’s own numbers show the gap: rebuilding a workflow that moved a meeting transcript from Google Drive to Salesforce this way cut the tokens it needed from 150,000 to 2,000, a 98.7% reduction, because the transcript never had to pass through the model’s context at all. The tradeoff is the sandbox itself. Code execution is Anthropic’s own pattern for calling tools over MCP specifically, and evaluating one of those tools for reliability now has to account for whatever the generated code gets wrong too, a failure mode ordinary tool calling doesn’t have.