# How do I grade a handoff between two agents?

Grading a handoff means checking two different objects, not one. The handoff span itself is thin: it records only the source and target agent names plus standard timing, so it tells you a handoff happened and how long it took but nothing about whether it happened correctly. The actual data lives in the run's items instead: a `HandoffCallItem` carries the model-generated arguments when the handoff declares an `input_type`, small structured metadata like a reason or priority, and by default the target agent still receives the entire prior conversation unless an `input_filter` trims it.

Two checks catch most real handoff failures. First, whether the `input_type` payload the model generated actually matches what the receiving agent needed, since a malformed or missing one raises before the handoff completes rather than passing through silently. Second, whether the model requested more than one handoff in the same turn: the SDK only executes the first and marks the span with an error, the same kind of new-path signal [Tessary's behavior_drift classifier](/answers/tessary-behavior-drift/what-is-tessarys-behavior-drift-classifier) watches for once a call site has an established baseline.

---

Sources:
- OpenAI Agents SDK docs: Handoffs: https://openai.github.io/openai-agents-python/handoffs/ (fetched 2026-09-18)

Source: https://tessary.ai/answers/openai-agents-sdk-evals/how-do-i-grade-a-handoff-between-two-agents
More on Openai agents sdk evals: https://tessary.ai/answers/openai-agents-sdk-evals
From Tessary, agent reliability for AI agents in production: https://tessary.ai
