# Agents API & Agents SDK Cheat Sheet

The three OpenAI agent runtimes compared, Agents API session lifecycle in Python and TypeScript, environments, tools, multi-agent, guardrails, tracing, billing and residency constraints.

import { Tabs, TabItem, Steps } from '@prosefly/astro-components';

OpenAI offers **three** ways to build agents, and knowing which to reach for is a recurring judgment on the developer tracks. This page compares them, then details the beta **Agents API** — OpenAI's managed Codex harness. Re-verify against [developers.openai.com/api/docs](https://developers.openai.com/api/docs); the Agents API is beta and its surface can change.

## The three runtimes

| Runtime | What it is | Choose it when |
| --- | --- | --- |
| **Responses API + tools** | You own the loop, model call by model call | Simple tool-using assistants, full control, no session or sandbox needs |
| **Agents SDK** | Open-source Python/TS framework you run: agent definitions, models/providers, running agents, sandboxed execution, orchestration, guardrails, results/state, integrations and observability, agent evals | You want code-first control, custom orchestration, self-hosted execution |
| **Agents API** (`/v1/agents/sessions`, header `OpenAI-Beta: agents=v1`, beta) | OpenAI-managed Codex harness: OpenAI runs sessions, orchestration, context compaction and recovery; you supply tools and pick the environment | Durable cloud agents, hosted sandboxes, long-running work, artifacts |

```text
   more control / more ops burden ◄─────────────────────► less ops / OpenAI-managed
   Responses API + tools        Agents SDK                 Agents API
   ─ your loop, your infra      ─ your infra, framework    ─ OpenAI runs the session
   ─ no sessions/sandbox        ─ orchestration+guardrails ─ hosted sandbox, compaction,
                                  you host                    recovery, subagents
```

## Agents API session lifecycle

The core concepts:

- **Agent** — model, instructions, tools, MCP servers.
- **Environment** — an OpenAI-hosted or self-hosted sandbox with files/artifacts, a lifecycle and security controls.
- **Session** — a durable instance: create → give it a task → follow progress via streaming or webhooks → continue or steer.
- **Events and items** — the stream of what the session did.

The managed harness provides sandboxed command/code execution, skills and instructions, MCP/tool data access, mid-run steering, context summarisation, subagent delegation and session resumption — you do not build any of that yourself.

<Tabs>
<TabItem label="Python">
```python
from openai import OpenAI

client = OpenAI()

session = client.beta.agents.sessions.create(
    extra_headers={"OpenAI-Beta": "agents=v1"},
    agent={
        "model": "gpt-6-astra",
        "instructions": "You are a build agent. Fix the failing tests, then stop.",
        "tools": [{"type": "shell"}, {"type": "apply_patch"}],
    },
    environment={"type": "hosted"},
    input="Clone the repo, run the test suite, and fix the first failing test.",
)
print(session.id, session.status)
```
</TabItem>
<TabItem label="TypeScript">
```typescript
import OpenAI from 'openai';

const client = new OpenAI();

const session = await client.beta.agents.sessions.create(
  {
    agent: {
      model: 'gpt-6-astra',
      instructions: 'You are a build agent. Fix the failing tests, then stop.',
      tools: [{ type: 'shell' }, { type: 'apply_patch' }],
    },
    environment: { type: 'hosted' },
    input: 'Clone the repo, run the test suite, and fix the first failing test.',
  },
  { headers: { 'OpenAI-Beta': 'agents=v1' } },
);
console.log(session.id, session.status);
```
</TabItem>
</Tabs>

The endpoint is `POST /v1/agents/sessions` and every call carries the `OpenAI-Beta: agents=v1` header. After creation you follow the session with streaming or webhooks, then continue it (add input) or steer it (redirect mid-run); the session is durable, so it survives beyond a single request and can be resumed.

## Environments

| Type | Where it runs | Notes |
| --- | --- | --- |
| **OpenAI-hosted sandbox** | OpenAI infrastructure | Fastest to start; container rates apply; files and artifacts persist for the session lifecycle |
| **Self-hosted sandbox** | Your infrastructure | You control the execution environment; **does not change the residency/ZDR constraints below** |

Both expose files and artifacts, a defined lifecycle (create → run → terminate) and security controls. The choice is about where code executes and what it can reach, not about data-handling guarantees.

## Tools

The agent's tools include everything on the built-in list (shell, code execution, file/web search, apply patch, image generation, computer use) plus:

- **MCP connections** — reach external systems through MCP servers; a **secure MCP tunnel** reaches private servers without exposing them.
- **Vaults** — managed secret storage the agent can use without the secret appearing in prompts or logs.
- **Skills and instructions** — reusable capability packages and standing guidance loaded into the session.

## Multi-agent

Delegate to subagents that run in parallel:

```json
{
  "multi_agent": { "enabled": true, "max_concurrent_subagents": 4 }
}
```

`max_concurrent_subagents` caps parallelism. More subagents finish fan-out work faster but multiply token spend and container cost, so size it to the task, not the maximum.

## Guardrails and approvals

- **Guardrails** — input/output checks that block or transform unsafe or off-policy content before it reaches the model or the user.
- **Approvals** — human-in-the-loop gates on sensitive actions; the session pauses for approval before executing (for example, before a shell command that mutates state).

Put approvals on side-effecting tools; put guardrails on the boundaries. The Agents SDK exposes both programmatically; the Agents API applies them within the managed harness.

## Tracing and observability

Sessions emit a trace of events and items — model calls, tool calls, subagent spans, steering interventions. Use it to debug why an agent did what it did and to build evals. The Agents SDK ships integrations and observability hooks; the Agents API exposes the event/item stream you follow during a session.

## Billing components

An Agents API session bills on three components:

| Component | What it covers |
| --- | --- |
| Model API rates | Per-token input/output for the model you chose |
| Standard tool rates | Per built-in tool invocation |
| Container rates | Compute for OpenAI-hosted sandboxes |

A self-hosted sandbox removes the container rate but you pay for your own compute instead — and it still does not relax the constraints below.

## Constraints

:::caution[US residency, no ZDR]
The Agents API is **US data residency only** and offers **no Zero Data Retention**. Choosing a self-hosted sandbox **does not** change either fact. A scenario that requires EU residency or ZDR for the data an agent will touch cannot be served by the Agents API — that is the discriminator to watch for.
:::

## Agents SDK surface

The open-source Agents SDK (Python/TS) covers, roughly in order:

<Steps>
1. **Agent definitions** — model, instructions, tools, handoffs.
2. **Models and providers** — which model backs the agent, including non-OpenAI providers.
3. **Running agents** — the run loop and results.
4. **Sandboxed execution** — run tools/code in a sandbox you control.
5. **Orchestration** — multi-agent handoffs and custom control flow.
6. **Guardrails** — input/output validation.
7. **Results and state** — capturing outputs and carrying state across runs.
8. **Integrations and observability** — tracing, logging, third-party hooks.
9. **Agent evals** — evaluate agent behaviour systematically.
</Steps>

## Choosing between them — worked reasoning

A team wants a durable coding agent that fixes CI failures overnight, keeps artifacts, and needs no on-call ops.

<Steps>
1. **Durable + long-running + artifacts + no ops** → this is the Agents API's home ground; the managed harness handles compaction, recovery and resumption.
2. **Check residency/ZDR** → the code and logs are internal but not regulated personal data, and the org accepts US residency and no ZDR. If either were required, the Agents API would be ruled out and the answer would be the **Agents SDK** on self-hosted infra.
3. **Model** → `gpt-6-astra` for the hardest fixes, or Terra/Luna for routine ones; size effort down where possible.
4. **Safety** → approvals on any tool that pushes or deploys; guardrails on inputs.
5. **Parallelism** → enable multi-agent with a modest `max_concurrent_subagents` and measure cost before raising it.
</Steps>

:::tip[Assessment signal]
"We don't want to host anything / OpenAI runs the session / durable long-running agent" → **Agents API**. "Data must stay in the EU" or "we need ZDR" → **not** the Agents API, even with a self-hosted sandbox → **Agents SDK**. "Simple tool-using assistant, full control, no sessions" → **Responses API + tools**.
:::

## Key facts to memorise

- Three runtimes: **Responses API + tools** (your loop), **Agents SDK** (your infra, framework), **Agents API** (OpenAI-managed harness, beta).
- Agents API: `POST /v1/agents/sessions`, header `OpenAI-Beta: agents=v1`, concepts Agent · Environment · Session · Events/items.
- Environments are OpenAI-hosted or self-hosted sandboxes; self-hosting does **not** change residency or ZDR.
- Multi-agent via `max_concurrent_subagents`; vaults for secrets; secure MCP tunnel for private servers.
- Billing = model rates + tool rates + container rates.
- **US residency only, no ZDR** — the hard constraint that rules the Agents API out for regulated data.
