AI Cert Prep
Type to search documentation.

Appendix · OpenAI

Model Lineup & Pricing (OpenAI)

OpenAI's September 2026 API model line-up, IDs and aliases, reasoning efforts, prices, context windows, specialised model families, a selection table and worked cost calculations.

This is the single reference for OpenAI model IDs, prices and limits used across the OpenAI tracks. Everything here reflects the published API line-up as of 15 September 2026; prices are USD per million tokens (MTok) and are rollout-sensitive, so treat the table as a snapshot to reason with, not a live rate card. Re-verify against the API pricing and API models pages before you commit an architecture or a budget.

API line-up (September 2026)

All latest models take text and image input, emit text output, are multilingual and support vision. They are served through the Responses API and the client SDKs.

ModelIDReasoning effortsInput $/MTokOutput $/MTokContextMax outputKnowledge cutoff
GPT-6 Astragpt-6-astralow, medium, high, xhigh, max$10$501.05M128K30 Apr 2026
GPT-5.6 Solgpt-5.6-sol (alias gpt-5.6)none…max$4$201.05M128K16 Feb 2026
GPT-5.6 Terragpt-5.6-terranone…max$2$121.05M128K16 Feb 2026
GPT-5.6 Lunagpt-5.6-lunanone…max$0.20$1.201.05M128K16 Feb 2026

Previous generation, still available: gpt-5.5 at $5 / $30 and gpt-5.5-pro at $30 / $180. Legacy models remain callable; new work should target the GPT-5.6 line or Astra.

Aliases move, pinned IDs do not

gpt-5.6 is an alias that resolves to gpt-5.6-sol. Aliases can be repointed by OpenAI; pin the explicit ID (gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-6-astra) in production so a server-side alias change never silently alters your cost or behaviour.

Reasoning effort

GPT-5.6 models accept efforts from none through max; Astra runs low, medium, high, xhigh, max. The rule straight from the docs is to use the lowest reasoning effort that gets the result — higher effort spends more output tokens (and therefore more money and latency) on internal reasoning. There is no exact mapping from GPT-5.5 efforts to GPT-5.6, so re-tune rather than porting an effort setting across a generation.

EffortSpend it on
none (5.6 only)Deterministic transforms where reasoning adds nothing
lowMechanical, well-specified tasks
mediumEveryday reasoning; a sensible default
highGenuinely hard, open-ended problems
xhighThe hardest multi-step work; Astra and 5.6
maxMaximum deliberation on one task; use sparingly

Model positioning

Selection guidance, straight from the docs:

Signal in the scenarioChooseWhy
Hardest end-to-end work: sustained reasoning, judgment, multi-tool orchestrationGPT-6 AstraThe frontier model; accept the $10 / $50 price
Complex, open-ended, high-value workGPT-5.6 SolStrong reasoning below Astra’s price
Pragmatic all-rounder; natural replacement for GPT-5.5 workloadsGPT-5.6 TerraBalanced cost and quality
Clear, repeatable, high-volume tasks: extraction, classification, transformation, structured summariesGPT-5.6 LunaCheapest by an order of magnitude
Migrating an existing GPT-5.5 workloadTerra first, benchmark, then decideClosest cost/quality match
text
capability ▲
│ GPT-6 Astra frontier, judgment, multi-tool
│ GPT-5.6 Sol complex, open-ended, high-value
│ GPT-5.6 Terra pragmatic all-rounder
│ GPT-5.6 Luna high-volume, structured, cheap
└───────────────────────────────► $ / MTok rises
choose the lowest tier that clears the task, then the lowest effort that clears it

ChatGPT-side model names

The consumer/business product exposes different labels for the same underlying family. Do not assume an API ID exists for a ChatGPT label or vice versa.

ChatGPT labelNotes
GPT-5.6 SolDefault reasoning-capable model
GPT-5.6 Sol ProHigher-tier Sol
GPT-5.6 TerraAll-rounder
GPT-5.6 LunaFast, lightweight
GPT-5 Thinking MiniLightweight thinking model
GPT-6 Pro (powered by GPT-6 Astra)On Pro $100 / Pro $200 / Business / Enterprise

Legacy models remain available in ChatGPT. The exact model roster a user sees is plan-dependent — see the ChatGPT feature and plan reference.

Specialised model families

These are purpose-built lines, several of them gated to authorised or approved organisations. Never route general traffic to them.

FamilyModelsPurpose / access
Cybergpt-5.6-cyber, gpt-daybreak-red-latest, gpt-daybreak-blue-latestAuthorised security work only
Life sciencesgpt-rosalind-researchLife-sciences research, approved orgs
Imagegpt-image-2.5-sunburst, gpt-image-2.5-flareImage generation
Realtime / GPT-Livegpt-live-1, gpt-realtime-2.1, gpt-realtime-2.1-mini, gpt-realtime-2, gpt-realtime-translate, gpt-realtime-1.5Low-latency voice / realtime and live translation
Transcriptiongpt-transcribe, gpt-live-transcribe, gpt-realtime-whisper, gpt-4o-transcribe, gpt-4o-mini-transcribeSpeech-to-text
TTSgpt-4o-mini-ttsText-to-speech

Gated families

The Daybreak (red/blue) and gpt-5.6-cyber models are for authorised security work only, and gpt-rosalind-research is limited to approved life-sciences organisations. A scenario that reaches for one of these for ordinary application traffic is choosing the wrong tool.

Selecting a model — decision table

The scenario says…ModelEffortDelivery
“extract fields from millions of documents, cheaply”Lunalow/noneBatch
“customer-facing assistant, balanced”TerramediumRealtime
“hard open-ended analysis, high value”SolhighRealtime
“the hardest agentic coding / research task we have”Astraxhigh/maxRealtime or Agents API
“offline nightly job, latency does not matter”any tieras neededBatch (discounted)
“the same long system prompt on every call”any tieras neededprompt caching
“we are migrating our GPT-5.5 workload”Terrabenchmarkcompare before switching

Worked cost calculations

Prices are per MTok. These illustrate order-of-magnitude differences; run your own numbers against live pricing before committing.

Example 1 — a single 50k-token RAG answer

A retrieval-augmented answer with 50,000 input tokens (retrieved context + prompt) and 1,000 output tokens, one call.

ModelInput costOutput costTotal
Luna50k × $0.20/1M = $0.0101k × $1.20/1M = $0.0012≈ $0.0112
Terra50k × $2/1M = $0.1001k × $12/1M = $0.012≈ $0.112
Sol50k × $4/1M = $0.2001k × $20/1M = $0.020≈ $0.220
Astra50k × $10/1M = $0.5001k × $50/1M = $0.050≈ $0.550

The input side dominates a RAG answer because retrieved context is large relative to the reply. Astra costs ~49× Luna for the same answer — so the model choice is a retrieval-quality decision, not a reasoning-power one, unless the synthesis is genuinely hard.

Example 2 — a 1M-token batch job

Classify or transform one million input tokens with 200k tokens of output, run offline.

ModelInputOutputRealtime totalWith Batch (offline)
Luna1M × $0.20 = $0.200.2M × $1.20 = $0.24$0.44cheapest; Batch lowers it further
Terra1M × $2 = $2.000.2M × $12 = $2.40$4.40reserve for harder items
Sol1M × $4 = $4.000.2M × $20 = $4.00$8.00rarely justified for classification

For structured high-volume work, Luna on Batch is the default; escalating a subset to Terra only where a validator flags low quality is a cascade, not a blanket upgrade.

Example 3 — an agent session

A single agent session that consumes 300k input tokens (accumulated context, tool results, compaction) and produces 40k output tokens across many turns.

ModelInputOutputTotal
Terra300k × $2/1M = $0.6040k × $12/1M = $0.48≈ $1.08
Sol300k × $4/1M = $1.2040k × $20/1M = $0.80≈ $2.00
Astra300k × $10/1M = $3.0040k × $50/1M = $2.00≈ $5.00

Agent sessions accumulate context, so compaction and context summarisation (see the Agents cheat sheet) are the real cost levers: a session that lets context grow unbounded on Astra is where budgets disappear. Prompt caching a stable system prompt across turns also helps.

Assessment signal

A stem that pairs “cheapest”, “high-volume”, “extraction” or “classification” with a latency-tolerant window is pointing at Luna on Batch. A stem that says “hardest”, “sustained reasoning” or “multi-tool judgment” points at Astra. “Natural replacement for our GPT-5.5 workload” is the Terra tell.

How to re-verify

  1. Open developers.openai.com/api/docs/models for the current IDs, aliases, context windows and cutoffs.
  2. Open developers.openai.com/api/docs/pricing for live input/output rates and any Batch/cached-input discounts.
  3. Check the Codex models page for the Codex-specific roster and retirement dates.
  4. When a fact here disagrees with the docs, the docs win — this page is a study aid, not an authority.

Key facts to memorise

  • Four current API models: Astra ($10/$50), Sol ($4/$20), Terra ($2/$12), Luna ($0.20/$1.20); all 1.05M context, 128K max output.
  • gpt-5.6 aliases to gpt-5.6-sol; pin explicit IDs in production.
  • Use the lowest tier and lowest effort that clears the task; there is no GPT-5.5 → 5.6 effort mapping.
  • Terra is the pragmatic replacement for GPT-5.5 workloads; Luna is for high-volume structured tasks.
  • Specialised families (cyber, Daybreak, Rosalind) are gated and off-limits for general traffic.

Last updated Sep 18, 2026