How do I grade a handoff between two agents?

Grading a handoff means checking two different objects, not one. The handoff span itself is thin: it records only the source and target agent names plus standard timing, so it tells you a handoff happened and how long it took but nothing about whether it happened correctly. The actual data lives in the run’s items instead: a HandoffCallItem carries the model-generated arguments when the handoff declares an input_type, small structured metadata like a reason or priority, and by default the target agent still receives the entire prior conversation unless an input_filter trims it.

Two checks catch most real handoff failures. First, whether the input_type payload the model generated actually matches what the receiving agent needed, since a malformed or missing one raises before the handoff completes rather than passing through silently. Second, whether the model requested more than one handoff in the same turn: the SDK only executes the first and marks the span with an error, the same kind of new-path signal Tessary’s behavior_drift classifier watches for once a call site has an established baseline.

sources

keep reading

More on this.

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y