Why did my agent call the wrong tool with invented arguments?

The model matched your request to whichever tool’s description sounded closest, then, having decided a tool was needed, filled in an argument it did not actually have with a plausible-looking guess instead of stopping to ask. Both failures trace to the same cause: tool calling gives the model no way to say it is unsure, only a schema to fill in, so a low-confidence match and a high-confidence one produce an identical-looking call. Vague or overlapping tool descriptions make the wrong-tool half worse, since the model is choosing on the text of the description, not the code behind it. Missing or ambiguous context earlier in the conversation makes the invented-argument half worse, since the model has a slot to fill and nothing real to put in it.

The fix lives in the tool definitions more than in the model: names and descriptions specific enough that no two tools plausibly match the same request, and required arguments the model can only get from context that is actually present.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y