What's the difference between tool calling and code execution?

Ordinary tool calling puts every call and its full result through the model’s own context, one round trip per tool: the model emits one call, waits, reads the result, and decides what to do next; code execution instead has the model write a short program, in a sandboxed runtime, that imports the available tools as functions and calls several of them directly, so only the code’s own output returns to the model’s context, not each tool’s raw response along the way.

Anthropic’s own numbers show the gap: rebuilding a workflow that moved a meeting transcript from Google Drive to Salesforce this way cut the tokens it needed from 150,000 to 2,000, a 98.7% reduction, because the transcript never had to pass through the model’s context at all. The tradeoff is the sandbox itself. Code execution is Anthropic’s own pattern for calling tools over MCP specifically, and evaluating one of those tools for reliability now has to account for whatever the generated code gets wrong too, a failure mode ordinary tool calling doesn’t have.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y