AI Cert Prep
Type to search documentation.

Appendix · OpenAI

OpenAI API Cheat Sheet

Responses API request and response shapes in Python and TypeScript, conversation state, streaming, background mode, structured outputs, tools, reasoning effort, caching, batch, errors and limits.

The Responses API is the primary interface for new work. Chat Completions is legacy — it still runs, but the Responses API is where conversation state, background mode, reasoning items and the built-in tool surface live, so build against it. Everything here reflects the September 2026 surface described in the OpenAI tracks; re-verify shapes against developers.openai.com/api/docs.

Minimal request

python
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY
resp = client.responses.create(
model="gpt-5.6-terra",
input="Give three risks of vendor lock-in.",
reasoning={"effort": "medium"},
)
print(resp.output_text)

input accepts a string or an array of typed items (messages, tool outputs, files). output_text is a convenience aggregate; the authoritative content is in the output array of items.

Request fields

FieldNotes
modelPinned ID: gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna
inputString or array of typed input items
instructionsSystem-level guidance for this response
reasoning.effortnone…max (5.6), low…max (Astra); lowest that works
max_output_tokensOutput cap; watch the incomplete status if hit
tools / tool_choiceBuilt-in tools and your functions
text.formatStructured output (JSON Schema)
previous_response_idServer-side conversation state
storePersist the response for later retrieval / state
streamtrue for SSE
backgroundtrue to run asynchronously
metadataYour key/value tags (e.g. hashed user id)

Response shape

json
{
"id": "resp_01",
"object": "response",
"model": "gpt-5.6-terra",
"status": "completed",
"output": [
{ "type": "reasoning", "id": "rs_01", "summary": [] },
{ "type": "message", "role": "assistant",
"content": [{ "type": "output_text", "text": "The notice period is 60 days." }] }
],
"usage": {
"input_tokens": 1200,
"input_tokens_details": { "cached_tokens": 1024 },
"output_tokens": 42,
"output_tokens_details": { "reasoning_tokens": 18 },
"total_tokens": 1242
}
}

status — branch on it

statusMeaningAction
completedFinished normallyRead output
incompleteStopped early (e.g. max_output_tokens)Inspect incomplete_details; continue or raise the cap
in_progressBackground/streamed, not donePoll or keep streaming
failedErroredInspect error; retry only if transient

output_tokens_details.reasoning_tokens are billed as output — high effort spends real money here.

Conversation state

Two ways to carry state; do not mix them for the same thread.

ApproachHowUse when
Server-sideSet store: true, then pass previous_response_id on the next callYou want OpenAI to hold the thread; less to send each turn
Client-sideResend the full input array yourselfYou need full control / your own store
python
first = client.responses.create(model="gpt-5.6-terra", input="My name is Dana.", store=True)
second = client.responses.create(
model="gpt-5.6-terra",
input="What is my name?",
previous_response_id=first.id,
)

Streaming

Server-sent events emit semantic events, not raw token deltas — branch on the event type.

python
stream = client.responses.create(
model="gpt-5.6-terra", input="Write a haiku about latency.", stream=True,
)
for event in stream:
if event.type == "response.output_text.delta":
print(event.delta, end="", flush=True)
elif event.type == "response.completed":
print("\n", event.response.usage)

Common event types: response.created, response.output_item.added, response.output_text.delta, response.function_call_arguments.delta, response.output_item.done, response.completed, response.error. Tool-call arguments stream as fragments — buffer and parse only on done.

Background mode

For long jobs, start with background: true, get an id back immediately, then poll or subscribe to a webhook.

python
job = client.responses.create(model="gpt-6-astra", input=big_task, background=True)
# later
resp = client.responses.retrieve(job.id)
while resp.status in ("queued", "in_progress"):
time.sleep(2)
resp = client.responses.retrieve(job.id)

Background mode pairs with webhooks so you are not holding an open connection for minutes. It is the API-level analogue of the Agents API’s durable sessions for one-shot long work.

Structured outputs

python
schema = {
"type": "object",
"properties": {
"vendor": {"type": "string"},
"total": {"type": "number"},
"currency": {"type": "string", "enum": ["USD", "EUR", "GBP"]},
},
"required": ["vendor", "total", "currency"],
"additionalProperties": False,
}
resp = client.responses.create(
model="gpt-5.6-terra",
input="Extract the invoice: Acme, 1240.50 USD.",
text={"format": {"type": "json_schema", "name": "invoice", "schema": schema, "strict": True}},
)

strict: true constrains generation to the schema. Still validate downstream and retry with the specific error fed back — structured output guarantees shape, not business correctness.

Function calling

python
tools = [{
"type": "function",
"name": "get_order",
"description": "Look up one order by ID. Returns status and ETA.",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"], "additionalProperties": False,
},
"strict": True,
}]
resp = client.responses.create(model="gpt-5.6-terra", input="Where is ORD-12345?", tools=tools)
# resp.output contains a function_call item; execute it, then send the output back:
followup = client.responses.create(
model="gpt-5.6-terra",
previous_response_id=resp.id,
input=[{"type": "function_call_output", "call_id": call_id,
"output": '{"status":"shipped","eta":"2026-09-17"}'}],
)

The loop: model emits a function_call item → you run it → you send a function_call_output item back (referencing previous_response_id or resending state) → repeat until a plain message.

Reasoning effort

python
client.responses.create(model="gpt-5.6-luna", input=simple_transform, reasoning={"effort": "none"})
client.responses.create(model="gpt-6-astra", input=hard_problem, reasoning={"effort": "xhigh"})

Use the lowest effort that gets the result. Reasoning tokens are billed as output. There is no exact GPT-5.5 → 5.6 effort mapping — re-tune per model. See the model lineup.

File inputs

python
f = client.files.create(file=open("contract.pdf", "rb"), purpose="user_data")
resp = client.responses.create(
model="gpt-5.6-terra",
input=[{"role": "user", "content": [
{"type": "input_file", "file_id": f.id},
{"type": "input_text", "text": "Summarise the termination clause."},
]}],
)

Images use input_image with a file_id or URL. Upload once and reference by id across many calls rather than re-uploading.

Compaction and token counting

  • Compaction summarises older turns server-side so a long session stays inside the context window while preserving the narrative. It is the API-side lever against unbounded context growth; the Agents API applies context summarisation automatically inside a session.
  • Token counting — inspect usage.input_tokens, usage.output_tokens, input_tokens_details.cached_tokens (cache hits) and output_tokens_details.reasoning_tokens (billed reasoning) on every response to keep cost honest.

Built-in tools

ToolOne-line purpose
web_searchAnswer from the live web with citations
file_searchRetrieve over your uploaded/indexed files
retrievalGrounded answers over a managed store
MCP / connectorsReach external systems via MCP servers
secure MCP tunnelReach private MCP servers without exposing them
code_interpreterRun code in a sandbox for data/analysis
image_generationGenerate images inline
computer_useDrive a computer/browser UI
shell / local shellExecute shell commands (sandboxed / local)
apply patchApply code edits
tool searchDiscover tools from a large catalogue
programmatic tool callingInvoke tools from generated code
async tool callingLong-running tools without blocking the turn

Quality and cost features

FeatureWhat it doesUse it when
Prompt cachingReuses a stable prefix; cached input is discounted; cache diagnostics report hitsThe same long system prompt/context repeats across calls
BatchOffline processing at a discount, results within a windowLatency-tolerant, high-volume jobs
Flex processingLower-priced, best-effort latency tierNon-urgent traffic that tolerates variable latency
Fast modeLatency-optimised pathInteractive, latency-critical calls
Predicted outputsSupply expected text to speed up editsRegenerating a document with small changes

Confirm cache hits via usage.input_tokens_details.cached_tokens; caching lowers cost more than it lowers rate-limit pressure.

Error codes and retry policy

HTTPTypeRetry?
400invalid_request_errorNo — fix the request
401authentication_errorNo — key/credentials
403permission_errorNo — entitlement/region
404not_found_errorNo — model/resource id
409conflictSometimes — resolve state then retry
422unprocessableNo — fix the payload
429rate_limit_errorYes — backoff, honour retry-after
500server_errorYes — backoff
503service_unavailableYes — backoff, consider a fallback model

Use exponential backoff with jitter, honour retry-after, log the request id from response headers, and pass an idempotency key on side-effecting requests so a retry does not double-act.

python
import time, random
from openai import OpenAI, RateLimitError, APIStatusError
client = OpenAI()
RETRYABLE = {429, 500, 503}
def call_with_retry(**params):
for attempt in range(6):
try:
return client.responses.create(**params)
except RateLimitError as e:
wait = float(e.response.headers.get("retry-after", 0)) or min(60, 2 ** attempt)
time.sleep(wait + random.uniform(0, 0.5))
except APIStatusError as e:
if e.status_code in RETRYABLE:
time.sleep(min(60, 2 ** attempt) + random.uniform(0, 0.5))
else:
raise
raise RuntimeError("exhausted retries")

Rate and spend limits

  • Rate limits bind on requests and tokens per minute; you hit whichever binds first. Watch the rate-limit response headers and throttle before a 429 rather than after.
  • Spend limits cap cost per period at the org/project level; hitting them returns an error, not a silent stop.
  • Manage both from the dashboard; enforce project-level budgets so one runaway job cannot exhaust the org.

Assessment signal

“Chat Completions” in a stem about new work is usually the distractor — the correct surface is the Responses API. “Long-running”, “don’t hold the connection”, “come back later” points at background mode; “same system prompt every call” points at prompt caching; “overnight, cheap” points at Batch.

Key facts to memorise

  • Responses API is primary; Chat Completions is legacy for new work.
  • Conversation state: store: true + previous_response_id (server-side) or resend input (client-side) — not both.
  • Streaming emits semantic events; branch on event type, buffer tool-arg fragments.
  • Structured output guarantees shape (strict: true), not business correctness — still validate and retry.
  • Retry only 429/5xx with backoff and retry-after; fix 4xx. Use idempotency keys on side-effecting calls.
  • Cached input, Batch, Flex, Fast mode and predicted outputs are the cost/latency levers.

Last updated Sep 18, 2026