All concepts

OpenAI Agents SDK

OpenAI's lightweight SDK for agents: an LLM loop with tools, handoffs, and guardrails.

Agentic AI · Intermediate · ~8 min

In plain English

A framework for building agents around one provider's models: tools, handoffs between agents, and guardrails, with the loop already written.

Why it's worth your time

It removes the boilerplate around the agent loop so you can spend your time on tools and evaluation instead.

If you remember three things

  • The loop, tool calling and handoffs come built in
  • Guardrails run as separate checks around the model
  • Framework choice is not architecture — the design rules are the same

Overview

The OpenAI Agents SDK is a small, production-focused Python framework for building agentic apps, and the successor to the experimental Swarm. An Agent is just an LLM configured with instructions and tools; a Runner executes the agent loop — calling the model, running any tool calls, and repeating until a final output. On top of that core it adds handoffs, guardrails, sessions, typed outputs, and built-in tracing, deliberately keeping the abstraction count low.

How it works

  1. Start: Agent An LLM configured with instructions and a set of tools.
  2. Agent -> Runner loop Calls the model, runs any tool calls, and repeats until a final output.
  3. Runner loop -> Tools Python functions whose JSON schemas are auto-generated from type hints.
  4. Tools -> Handoffs Delegate to a specialist sub-agent — the routing pattern from Swarm.
  5. Handoffs -> Guardrails Input/output validation that can trip a tripwire and halt the run.
  6. Guardrails -> Typed output A Pydantic-validated result, with tracing on every span.

In an interview

It's OpenAI's minimal SDK for building agents, the production successor to Swarm. You define an Agent as an LLM plus instructions and tools, and Runner.run drives the loop of calling the model and running tool calls until a final answer. It layers on handoffs to specialist sub-agents, input/output guardrails, automatic session history, Pydantic-typed outputs, and tracing — while staying provider-agnostic.

Production defaults

Same rules apply
5–8 tools, hard step caps, out-of-model approval for side effects
Tracing
turn it on from day one. Agent debugging without traces is guesswork
Portability
keep prompts, tools and evals in your own code so the framework stays replaceable

What breaks

  • Behaviour differs from your local tests — Model version drift. Pin the model and run evals on every change.
  • Locked to one provider — Your tools and prompts should live outside the framework. That's what makes a switch a week, not a rewrite.

Watch it explained

OpenAI just made your entire tech stack obsolete... — Fireship, 4:18

Related