Why do tool outputs need to be captured instead of called again live?

Tool outputs need to be captured, not re-fetched, because the systems behind them keep moving. A database row updates, an inventory count changes, a search index reindexes, so calling the same tool with the same arguments an hour later can return a different result than the one the agent actually saw. If replay re-calls tools live, you’re not replaying the failure anymore, you’re running a new session against current state that happens to share a starting point. The fix you’re testing might pass only because the data changed underneath it, not because the fix works. Recording each tool’s output at the time it was returned, and replaying it verbatim instead of the live call, is what keeps the replayed case identical to the one that actually failed.

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y