all answers

Frameworks and tooling

Claude agent SDK evals

The Claude Agent SDK is the harness behind Claude Code, published so you can build your own agent on the same loop: gather context, act, check the result, repeat. It gives you the tool loop, file and shell tools, subagents, MCP support, sessions you can resume or fork, and hooks at fixed points in the loop.

Hooks and subagents are the two features that matter for evaluation. A hook is a callback at a named point, such as before a tool call or after one. That is where a check or a trace emit attaches without touching the agent's prompt. A subagent runs in its own context and returns a result to the main loop, so it is a natural unit to grade on its own instead of through whatever the main agent finally says.

Telemetry is OpenTelemetry, turned on with CLAUDE_CODE_ENABLE_TELEMETRY and pointed with the standard OTEL exporter variables. Metrics and events are the settled part: tokens and cost per session, and an event for each prompt, response, API request, API error, tool result, and tool decision. Spans are still behind a beta flag, and their names are Claude Code's own rather than the OpenTelemetry GenAI ones, so anything reading them has to map them first.

17 questions

Answered, plainly.

How do I trace a Claude Agent SDK agent?Set CLAUDE_CODE_ENABLE_TELEMETRY=1 and an OTLP exporter per signal; traces also need CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1, since they're still in beta.answer →How do I evaluate a subagent separately from the main loop?Grade the subagent's own final message from the Agent tool result, not the parent's reply, since the parent may summarize or rephrase it before answering.answer →How do I check a Claude Agent SDK tool call without touching the prompt?Register a PreToolUse or PostToolUse hook: a callback the SDK invokes before or after a tool runs, entirely outside the model's context or its prompt.answer →What does Claude Code's OpenTelemetry output contain?Three signals: metrics for tokens, cost, and sessions, log events per prompt and tool result, and beta trace spans named claude_code.*, not gen_ai.*.answer →What can a Claude Agent SDK hook check?Far more than tool calls: session start and end, a prompt before the model sees it, compaction, subagent completion, model switches, and permission decisions.answer →What's the difference between the Claude Agent SDK and Claude Code?The SDK is the same loop, tools, and context management that runs Claude Code, packaged as a library; Claude Code is Anthropic's own product built on it.answer →How do I catch a regression in a Claude Agent SDK agent?There's no built-in regression feature. Build one from a hook that scores each run and a forked session that replays a fixed case against the changed agent.answer →What's the difference between the Claude Agent SDK and Google's Agent Development Kit (ADK)?The Claude Agent SDK runs only Claude models as Claude Code's own loop; Google's ADK is model-agnostic and built around named multi-agent workflow types.answer →What's the difference between the Claude Agent SDK and the OpenAI Agents SDK?The Claude Agent SDK runs only Claude models as Claude Code's own loop; the OpenAI Agents SDK defaults to OpenAI models and traces into its own backend, not OpenTelemetry.answer →What's the difference between the Claude Agent SDK and AWS Strands Agents?The Claude Agent SDK runs only Claude as Claude Code's own loop; AWS Strands is model-agnostic across Bedrock, OpenAI, and Gemini, with native OTel tracing and five multi-agent patterns.answer →What's the difference between the Claude Agent SDK and Claude Managed Agents?The Claude Agent SDK runs in your own process; Claude Managed Agents is Anthropic's hosted harness with server-side sessions, ineligible for zero data retention.answer →What's the difference between the Claude Agent SDK and Deep Agents?The Claude Agent SDK only runs Claude, embedding Claude Code's own loop; Deep Agents is a model-agnostic LangChain library on the LangGraph runtime, with the same planning and file-tool shape.answer →What's the difference between the Claude Agent SDK and LangGraph?The Claude Agent SDK only runs Claude in an implicit loop; LangGraph is a model-agnostic runtime you wire by hand, with per-step checkpointing built in.answer →What's the difference between the Claude Agent SDK and LangChain?The Claude Agent SDK wraps the Claude Code binary and only runs Claude; LangChain is model-agnostic, reaching Claude, GPT, Gemini, or any other provider through one API.answer →Can untrusted text trigger a file read in the Claude Agent SDK?Yes, by default: the SDK reads an @path reference or a leading slash command in any prompt text as a command, not just text a person typed.answer →Is the Claude Agent SDK open source?Only the Python package: claude-agent-sdk-python is MIT-licensed. The TypeScript package is all-rights-reserved under Anthropic's commercial terms.answer →What's the difference between the Claude Agent SDK and the Claude API?The Agent SDK ships a full agent loop built on Claude Code; the plain Claude API (Anthropic's Client SDK) gives raw access and leaves the loop to you.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y