all answers

Frameworks and tooling

Openai agents SDK evals

The OpenAI Agents SDK is a small runtime for building agents. An agent is a model with instructions and tools. A handoff passes the conversation to another agent. Guardrails run checks on input and output. Sessions carry history between turns. The loop around all of this is thin, which helps evaluation, because there's very little framework behaviour between what you wrote and what ran.

Tracing is built in and on by default. One run is one trace, and the SDK already splits it into the pieces you'd want to grade: a span for each agent, each model call, each tool call, each handoff, and each guardrail. Two of those are worth more than they look. A handoff span records the moment control moved from one agent to another and what went with it, which is where most multi-agent failures start. A guardrail span records that a check ran and what it decided. That is a separate fact from the model's output, and the rate at which guardrails trip is worth tracking on its own.

Getting the spans out of OpenAI's backend takes a trace processor. There is no OpenTelemetry exporter in the SDK. add_trace_processor sends spans to a second destination alongside OpenAI's, set_trace_processors replaces it, and OPENAI_AGENTS_DISABLE_TRACING turns tracing off entirely.

16 questions

Answered, plainly.

How do I trace an OpenAI Agents SDK run?Every OpenAI Agents SDK run traces by default, with a span per agent, model call, tool call, handoff, and guardrail, sent to OpenAI's dashboard unless redirected.answer →What happens when an OpenAI Agents SDK guardrail trips?An agent guardrail tripwire stops the run and raises an exception. A tool guardrail can reject just the tool call or stop the run, with a different result shape.answer →Should a guardrail result be graded separately from the output?Yes: a guardrail is checking a different thing than the output is, so it needs its own pass rate and its own dataset of cases that should and shouldn't trip.answer →Do OpenAI Agents SDK guardrails run on every agent after a handoff?No. Input guardrails only check the first agent and output guardrails only the last, so an agent picked up mid-handoff runs with no guardrail of its own.answer →How do I grade a handoff between two agents?Check two objects, not one: the handoff span only records the source and target agent names, while the actual transferred data lives in the run's items.answer →What's the difference between the OpenAI Agents SDK and LangGraph?LangGraph is a graph of nodes with a checkpointed state you can rewind to any step; the Agents SDK is agents passing a conversation through handoffs.answer →How do I catch a regression in an Agents SDK workflow?Grade production traces first to see what a change broke, then promote the cases that matter into a fixed dataset you rerun and compare after every change.answer →What's the difference between the OpenAI Agents SDK and LangChain?The Agents SDK is a few composable primitives you wire directly; LangChain is a bigger harness whose own agents now run on LangGraph underneath.answer →What's the difference between the OpenAI Agents SDK and Google's Agent Development Kit (ADK)?The OpenAI Agents SDK defaults to OpenAI's own models through its Responses API; Google's ADK is model-agnostic and organizes agents into named workflow types.answer →What's the difference between the OpenAI Agents SDK and the Responses API?The Responses API is one call to a model with no orchestration; the Agents SDK is a runtime built on top of it, adding the loop, handoffs, guardrails, and sessions.answer →What's the difference between the OpenAI Agents SDK and CrewAI?The OpenAI Agents SDK is a fixed set of primitives with tracing on by default; CrewAI splits into autonomous Crews and a Flows layer, with tracing that needs an AMP account and opt-in.answer →What's the difference between the OpenAI Agents SDK and Pydantic AI?The OpenAI Agents SDK defaults to OpenAI's models with proprietary tracing; Pydantic AI is model-agnostic, validates every call with Pydantic types, and emits OpenTelemetry.answer →What's the difference between the OpenAI Agents SDK and AutoGen?The Agents SDK moves work by handoff, one agent delegating to a named specialist; AutoGen's signature pattern is a group chat, with a manager choosing who speaks next.answer →What's the difference between the OpenAI Agents SDK and the Codex SDK?The Agents SDK is a general framework for building any agent from a few primitives; the Codex SDK only starts, resumes, and streams threads on OpenAI's own Codex harness.answer →Does the OpenAI Agents SDK support MCP tools?Yes. The SDK ships dedicated MCP server classes, MCPServerStdio, MCPServerStreamableHttp, and HostedMCPTool, that an agent uses like any other tool.answer →Is the OpenAI Agents SDK open source?Yes. Both openai-agents-python and openai-agents-js ship under the MIT license, so the SDK itself is free to read, modify, and redistribute.answer →

Two ways to run Tessary.

Tessary is an open-source agent reliability platform. Cloud and self-hosted run the same workflow on the OpenTelemetry traces your agent already emits.

Tessary Cloud

We host it for you. Send your first trace with nothing to deploy and no model key.

what's includedper organization
traces
10,000 per calendar month
stored trace data
1 GB
retention
30 days
model credit
$10, one-time, for triage and root-cause analysis
credit card
not required

Self-hosted Tessary

Run the open-source code on your own infrastructure with one command. Add your own model key for triage and root-cause analysis.

Self-host Tessary for me by following https://github.com/tessaryai/tessary/blob/main/setup.md

docker compose -f oci://docker.io/tessaryai/tessary:compose up -d -y