Appendix · OpenAI
OpenAI API Cheat Sheet
Responses API request and response shapes in Python and TypeScript, conversation state, streaming, background mode, structured outputs, tools, reasoning effort, caching, batch, errors and limits.
The Responses API is the primary interface for new work. Chat Completions is legacy — it still runs, but the Responses API is where conversation state, background mode, reasoning items and the built-in tool surface live, so build against it. Everything here reflects the September 2026 surface described in the OpenAI tracks; re-verify shapes against developers.openai.com/api/docs.
Minimal request
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY
resp = client.responses.create( model="gpt-5.6-terra", input="Give three risks of vendor lock-in.", reasoning={"effort": "medium"},)print(resp.output_text)import OpenAI from 'openai';
const client = new OpenAI(); // reads OPENAI_API_KEY
const resp = await client.responses.create({ model: 'gpt-5.6-terra', input: 'Give three risks of vendor lock-in.', reasoning: { effort: 'medium' },});console.log(resp.output_text);input accepts a string or an array of typed items (messages, tool outputs, files). output_text is a convenience aggregate; the authoritative content is in the output array of items.
Request fields
| Field | Notes |
|---|---|
model | Pinned ID: gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna |
input | String or array of typed input items |
instructions | System-level guidance for this response |
reasoning.effort | none…max (5.6), low…max (Astra); lowest that works |
max_output_tokens | Output cap; watch the incomplete status if hit |
tools / tool_choice | Built-in tools and your functions |
text.format | Structured output (JSON Schema) |
previous_response_id | Server-side conversation state |
store | Persist the response for later retrieval / state |
stream | true for SSE |
background | true to run asynchronously |
metadata | Your key/value tags (e.g. hashed user id) |
Response shape
{ "id": "resp_01", "object": "response", "model": "gpt-5.6-terra", "status": "completed", "output": [ { "type": "reasoning", "id": "rs_01", "summary": [] }, { "type": "message", "role": "assistant", "content": [{ "type": "output_text", "text": "The notice period is 60 days." }] } ], "usage": { "input_tokens": 1200, "input_tokens_details": { "cached_tokens": 1024 }, "output_tokens": 42, "output_tokens_details": { "reasoning_tokens": 18 }, "total_tokens": 1242 }}status — branch on it
status | Meaning | Action |
|---|---|---|
completed | Finished normally | Read output |
incomplete | Stopped early (e.g. max_output_tokens) | Inspect incomplete_details; continue or raise the cap |
in_progress | Background/streamed, not done | Poll or keep streaming |
failed | Errored | Inspect error; retry only if transient |
output_tokens_details.reasoning_tokens are billed as output — high effort spends real money here.
Conversation state
Two ways to carry state; do not mix them for the same thread.
| Approach | How | Use when |
|---|---|---|
| Server-side | Set store: true, then pass previous_response_id on the next call | You want OpenAI to hold the thread; less to send each turn |
| Client-side | Resend the full input array yourself | You need full control / your own store |
first = client.responses.create(model="gpt-5.6-terra", input="My name is Dana.", store=True)second = client.responses.create( model="gpt-5.6-terra", input="What is my name?", previous_response_id=first.id,)Streaming
Server-sent events emit semantic events, not raw token deltas — branch on the event type.
stream = client.responses.create( model="gpt-5.6-terra", input="Write a haiku about latency.", stream=True,)for event in stream: if event.type == "response.output_text.delta": print(event.delta, end="", flush=True) elif event.type == "response.completed": print("\n", event.response.usage)const stream = await client.responses.create({ model: 'gpt-5.6-terra', input: 'Write a haiku about latency.', stream: true,});for await (const event of stream) { if (event.type === 'response.output_text.delta') process.stdout.write(event.delta); else if (event.type === 'response.completed') console.log('\n', event.response.usage);}Common event types: response.created, response.output_item.added, response.output_text.delta, response.function_call_arguments.delta, response.output_item.done, response.completed, response.error. Tool-call arguments stream as fragments — buffer and parse only on done.
Background mode
For long jobs, start with background: true, get an id back immediately, then poll or subscribe to a webhook.
job = client.responses.create(model="gpt-6-astra", input=big_task, background=True)# laterresp = client.responses.retrieve(job.id)while resp.status in ("queued", "in_progress"): time.sleep(2) resp = client.responses.retrieve(job.id)Background mode pairs with webhooks so you are not holding an open connection for minutes. It is the API-level analogue of the Agents API’s durable sessions for one-shot long work.
Structured outputs
schema = { "type": "object", "properties": { "vendor": {"type": "string"}, "total": {"type": "number"}, "currency": {"type": "string", "enum": ["USD", "EUR", "GBP"]}, }, "required": ["vendor", "total", "currency"], "additionalProperties": False,}resp = client.responses.create( model="gpt-5.6-terra", input="Extract the invoice: Acme, 1240.50 USD.", text={"format": {"type": "json_schema", "name": "invoice", "schema": schema, "strict": True}},)strict: true constrains generation to the schema. Still validate downstream and retry with the specific error fed back — structured output guarantees shape, not business correctness.
Function calling
tools = [{ "type": "function", "name": "get_order", "description": "Look up one order by ID. Returns status and ETA.", "parameters": { "type": "object", "properties": {"order_id": {"type": "string"}}, "required": ["order_id"], "additionalProperties": False, }, "strict": True,}]
resp = client.responses.create(model="gpt-5.6-terra", input="Where is ORD-12345?", tools=tools)# resp.output contains a function_call item; execute it, then send the output back:followup = client.responses.create( model="gpt-5.6-terra", previous_response_id=resp.id, input=[{"type": "function_call_output", "call_id": call_id, "output": '{"status":"shipped","eta":"2026-09-17"}'}],)The loop: model emits a function_call item → you run it → you send a function_call_output item back (referencing previous_response_id or resending state) → repeat until a plain message.
Reasoning effort
client.responses.create(model="gpt-5.6-luna", input=simple_transform, reasoning={"effort": "none"})client.responses.create(model="gpt-6-astra", input=hard_problem, reasoning={"effort": "xhigh"})Use the lowest effort that gets the result. Reasoning tokens are billed as output. There is no exact GPT-5.5 → 5.6 effort mapping — re-tune per model. See the model lineup.
File inputs
f = client.files.create(file=open("contract.pdf", "rb"), purpose="user_data")resp = client.responses.create( model="gpt-5.6-terra", input=[{"role": "user", "content": [ {"type": "input_file", "file_id": f.id}, {"type": "input_text", "text": "Summarise the termination clause."}, ]}],)Images use input_image with a file_id or URL. Upload once and reference by id across many calls rather than re-uploading.
Compaction and token counting
- Compaction summarises older turns server-side so a long session stays inside the context window while preserving the narrative. It is the API-side lever against unbounded context growth; the Agents API applies context summarisation automatically inside a session.
- Token counting — inspect
usage.input_tokens,usage.output_tokens,input_tokens_details.cached_tokens(cache hits) andoutput_tokens_details.reasoning_tokens(billed reasoning) on every response to keep cost honest.
Built-in tools
| Tool | One-line purpose |
|---|---|
web_search | Answer from the live web with citations |
file_search | Retrieve over your uploaded/indexed files |
| retrieval | Grounded answers over a managed store |
| MCP / connectors | Reach external systems via MCP servers |
| secure MCP tunnel | Reach private MCP servers without exposing them |
code_interpreter | Run code in a sandbox for data/analysis |
image_generation | Generate images inline |
computer_use | Drive a computer/browser UI |
| shell / local shell | Execute shell commands (sandboxed / local) |
| apply patch | Apply code edits |
| tool search | Discover tools from a large catalogue |
| programmatic tool calling | Invoke tools from generated code |
| async tool calling | Long-running tools without blocking the turn |
Quality and cost features
| Feature | What it does | Use it when |
|---|---|---|
| Prompt caching | Reuses a stable prefix; cached input is discounted; cache diagnostics report hits | The same long system prompt/context repeats across calls |
| Batch | Offline processing at a discount, results within a window | Latency-tolerant, high-volume jobs |
| Flex processing | Lower-priced, best-effort latency tier | Non-urgent traffic that tolerates variable latency |
| Fast mode | Latency-optimised path | Interactive, latency-critical calls |
| Predicted outputs | Supply expected text to speed up edits | Regenerating a document with small changes |
Confirm cache hits via usage.input_tokens_details.cached_tokens; caching lowers cost more than it lowers rate-limit pressure.
Error codes and retry policy
| HTTP | Type | Retry? |
|---|---|---|
| 400 | invalid_request_error | No — fix the request |
| 401 | authentication_error | No — key/credentials |
| 403 | permission_error | No — entitlement/region |
| 404 | not_found_error | No — model/resource id |
| 409 | conflict | Sometimes — resolve state then retry |
| 422 | unprocessable | No — fix the payload |
| 429 | rate_limit_error | Yes — backoff, honour retry-after |
| 500 | server_error | Yes — backoff |
| 503 | service_unavailable | Yes — backoff, consider a fallback model |
Use exponential backoff with jitter, honour retry-after, log the request id from response headers, and pass an idempotency key on side-effecting requests so a retry does not double-act.
import time, randomfrom openai import OpenAI, RateLimitError, APIStatusError
client = OpenAI()RETRYABLE = {429, 500, 503}
def call_with_retry(**params): for attempt in range(6): try: return client.responses.create(**params) except RateLimitError as e: wait = float(e.response.headers.get("retry-after", 0)) or min(60, 2 ** attempt) time.sleep(wait + random.uniform(0, 0.5)) except APIStatusError as e: if e.status_code in RETRYABLE: time.sleep(min(60, 2 ** attempt) + random.uniform(0, 0.5)) else: raise raise RuntimeError("exhausted retries")Rate and spend limits
- Rate limits bind on requests and tokens per minute; you hit whichever binds first. Watch the rate-limit response headers and throttle before a 429 rather than after.
- Spend limits cap cost per period at the org/project level; hitting them returns an error, not a silent stop.
- Manage both from the dashboard; enforce project-level budgets so one runaway job cannot exhaust the org.
Assessment signal
“Chat Completions” in a stem about new work is usually the distractor — the correct surface is the Responses API. “Long-running”, “don’t hold the connection”, “come back later” points at background mode; “same system prompt every call” points at prompt caching; “overnight, cheap” points at Batch.
Key facts to memorise
- Responses API is primary; Chat Completions is legacy for new work.
- Conversation state:
store: true+previous_response_id(server-side) or resendinput(client-side) — not both. - Streaming emits semantic events; branch on event
type, buffer tool-arg fragments. - Structured output guarantees shape (
strict: true), not business correctness — still validate and retry. - Retry only 429/5xx with backoff and
retry-after; fix 4xx. Use idempotency keys on side-effecting calls. - Cached input, Batch, Flex, Fast mode and predicted outputs are the cost/latency levers.
Last updated Sep 18, 2026