How do I evaluate a subagent separately from the main loop?
Grade the subagent’s own final message, carried back to the parent as the Agent tool’s result, rather than whatever the parent agent says afterward. A subagent’s intermediate tool calls and reasoning stay inside its own isolated context and never reach the parent; only that one final message crosses the boundary, and the parent is free to summarize or rephrase it before it reaches the user, so a check run on the parent’s reply is really grading the parent’s paraphrase, not the subagent’s actual work.
Messages carry a parent_tool_use_id field that ties them back to the subagent that produced them, which is what lets a grader pull just that subagent’s output out of the stream instead of the whole session. A subagent’s transcript is also stored separately from the parent’s and can be resumed on its own, so a failing case can be replayed at the point the subagent actually went wrong. Grading one node of a LangGraph graph is the same idea applied to a different framework’s unit of delegation.