How do I grade a handoff between two agents?
Grading a handoff means checking two different objects, not one. The handoff span itself is thin: it records only the source and target agent names plus standard timing, so it tells you a handoff happened and how long it took but nothing about whether it happened correctly. The actual data lives in the run’s items instead: a HandoffCallItem carries the model-generated arguments when the handoff declares an input_type, small structured metadata like a reason or priority, and by default the target agent still receives the entire prior conversation unless an input_filter trims it.
Two checks catch most real handoff failures. First, whether the input_type payload the model generated actually matches what the receiving agent needed, since a malformed or missing one raises before the handoff completes rather than passing through silently. Second, whether the model requested more than one handoff in the same turn: the SDK only executes the first and marks the span with an error, the same kind of new-path signal Tessary’s behavior_drift classifier watches for once a call site has an established baseline.
sources
- OpenAI Agents SDK docs: Handoffs fetched