Appendix · Claude
Model Lineup & Pricing
Current Claude models, IDs, limits, prices, breaking changes and worked cost calculations as of September 2026.
Current lineup (September 2026)
| Model | ID | Context | Max output | Input / Output per MTok | Cache read | Positioning |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | claude-fable-5-1 | 1M | 128k | $10 / $50 | $0.25 | Most capable; hardest reasoning and agentic work |
| Claude Opus 5 | claude-opus-5 | 1M | 128k | $5 / $25 | $0.50 | Default for complex agentic coding and enterprise |
| Claude Sonnet 5 | claude-sonnet-5 | 1M | 128k | $2 / $10 | $0.20 | Best speed / intelligence balance |
| Claude Haiku 4.5 | claude-haiku-4-5 | 200k | 64k | $1 / $5 | $0.10 | Fastest and cheapest; classification, routing, extraction |
Cache writes cost ≈ 1.25× base input (5-minute TTL) or 2× (1-hour TTL). Batch API: 50% off input and output.
Legacy, still available: Fable 5, Opus 4.8, 4.7, 4.6, 4.5, Sonnet 4.6, 4.5. Retired: Opus 4.1 (2026-08-05).
Published retirement floors: Fable 5.1 not before 2027-09-01 · Opus 5 not before 2027-07-24 · Sonnet 5 not before 2027-06-30 · Haiku 4.5 not before 2026-10-15.
Verify live
Model IDs from the 4.6 generation onward are dateless but still pinned snapshots. Query client.models.list() and client.models.retrieve(id) for live max_input_tokens, max_tokens and capabilities rather than trusting any static table – including this one.
Capability matrix
| Capability | Fable 5.1 | Opus 5 | Sonnet 5 | Haiku 4.5 |
|---|---|---|---|---|
Adaptive thinking {"type":"adaptive"} | ✓ (always on) | ✓ | ✓ | – |
budget_tokens thinking | 400 error | 400 error | 400 error | ✓ (only model) |
effort low/medium/high/xhigh | ✓ | ✓ | ✓ (no xhigh benefit) | – |
tool_choice: "any" / forced tool | 400 error | ✓ | ✓ | ✓ |
Structured outputs output_config.format | ✓ | ✓ | ✓ | ✓ |
strict: true tool schemas | ✓ | ✓ | ✓ | ✓ |
Mid-conversation role: "system" messages | ✓ | ✓ | – | – |
| Task budgets | ✓ | ✓ | – | – |
| Thinking blocks portable to other models | – (bound) | ✓ | ✓ | n/a |
| Zero Data Retention | – (30-day required) | ✓ | ✓ | ✓ |
| Priority Tier | – | – | – | ✓ |
| Prompt caching | ✓ | ✓ | ✓ | ✓ (2048-token minimum) |
| Batch API | ✓ | ✓ | ✓ | ✓ |
Fable 5.1 breaking changes (exam favourites)
- No forced tool use.
tool_choice: "any"and{"type": "tool", "name": …}return 400. Alternatives:autoplus an explicit instruction;strict: trueon the tool schema;output_config.formatstructured output. - Thinking-block binding. Thinking blocks are readable only by the model that produced them or a newer one. A fallback or router switch to an older model silently drops them (unbilled) and the older model re-plans from scratch.
- Append-only history. Editing, reordering or removing earlier turns invalidates every later thinking block. Freeze
systemandtools, deliver mid-session instruction changes asrole: "system"messages, and trim context server-side (context editing / compaction) rather than on the client. - Retention. Requires 30-day data retention; not available to ZDR organisations unless authorised. Excluded from Priority Tier (as are Opus 5 and Sonnet 5).
Selecting a model – decision table
| Signal in the scenario | Choose | Why |
|---|---|---|
| “Classify”, “route”, “extract simple fields”, “millions of items”, “sub-second” | Haiku 4.5 | Cheapest, fastest; quality sufficient for narrow tasks |
| “General assistant”, “customer-facing chat”, “balanced cost and quality” | Sonnet 5 | Best trade-off; 1M context |
| “Complex agentic coding”, “multi-step reasoning”, “enterprise default” | Opus 5 | Default for hard agentic work |
| “Hardest research”, “frontier reasoning”, budget is secondary | Fable 5.1 | Most capable; accept breaking-change constraints |
| “Cost matters, latency does not” | Any tier + Batch API | 50% discount |
| Same long prefix reused | Any tier + prompt caching | ~90% saving on cached reads |
| Mixed difficulty stream | Cascade: Haiku → Sonnet → Opus | Cheap model handles the bulk; escalate on low confidence from a validator, not self-report |
Worked cost calculations
Example 1 – 10,000 documents, 3,000 input tokens each, 500 output tokens each
| Approach | Input cost | Output cost | Total |
|---|---|---|---|
| Haiku 4.5 realtime | 30M × $1 = $30 | 5M × $5 = $25 | $55 |
| Sonnet 5 realtime | 30M × $2 = $60 | 5M × $10 = $50 | $110 |
| Opus 5 realtime | 30M × $5 = $150 | 5M × $25 = $125 | $275 |
| Sonnet 5 Batch (50%) | $30 | $25 | $55 |
| Opus 5 Batch (50%) | $75 | $62.50 | $137.50 |
Example 2 – prompt caching on a 20,000-token system prompt + tools, 1,000 requests/day on Sonnet 5
| Without caching | With caching (5-min TTL, ~1 write per 5 min ≈ 288 writes/day) | |
|---|---|---|
| Prefix tokens billed | 20M at $2 = $40/day | Writes: 288 × 20k × $2.50 = $14.40; Reads: 712 × 20k × $0.20 = $2.85 → ≈ $17.25/day |
| Saving | – | ≈ 57% on the prefix; approaches 90% as request rate rises |
Rule: caching pays back after roughly two reads of the same prefix within the TTL.
Example 3 – choosing effort
| Task | Effort | Rationale |
|---|---|---|
| Subagent that renames files per a spec | low | Mechanical |
| Summarise a 50-page contract | medium | Moderate reasoning, cost-sensitive |
| Default production reasoning | high | Baseline |
| Multi-hour refactor across 40 files (Opus 5 / Fable 5.1) | xhigh | Hardest agentic work |
Cloud availability
| Platform | Notes |
|---|---|
| Claude API (direct) | Full feature surface first; simplest |
| Amazon Bedrock | AWS IAM, data residency, FedRAMP High, existing AWS spend |
| Google Vertex AI | GCP IAM, residency, FedRAMP High |
| Microsoft Foundry | Azure ecosystem |
Feature availability can lag on cloud platforms; check per-feature docs before committing an architecture.
Feature availability by cloud platform
The exam tests the reflex that the direct Claude API gets new features first, and that a governance or residency requirement can force a platform that lags on a feature you depend on. Treat this as a snapshot to reason with, not a live SLA.
| Feature | Claude API | Bedrock | Vertex AI | Foundry |
|---|---|---|---|---|
| Newest model on launch day | ✓ first | Usually days–weeks later | Usually days–weeks later | Varies |
| Prompt caching | ✓ | ✓ | ✓ | Check |
| Message Batches | ✓ | ✓ (batch inference) | ✓ (batch prediction) | Check |
| MCP connector (server-side) | ✓ | Lags | Lags | Lags |
| Files API / citations | ✓ | Partial | Partial | Check |
Structured outputs output_config.format | ✓ | ✓ | ✓ | Check |
| Context editing / compaction | ✓ | Lags | Lags | Lags |
| FedRAMP High authorisation | – | ✓ | ✓ | – |
| Data residency controls | Limited | ✓ (region) | ✓ (region) | ✓ (region) |
| IAM / enterprise identity | API keys | AWS IAM | GCP IAM | Entra ID |
Exam signal
“We are on AWS with FedRAMP High and IAM already” → Bedrock, even if it means waiting for a feature. “We need the MCP connector and context editing today” → direct Claude API. A stem that pairs a residency requirement with a bleeding-edge feature is testing whether you notice the platform lag.
Rate-limit tiers and dimensions
Rate limits apply on three dimensions simultaneously; you hit whichever binds first.
| Dimension | Meaning | Typical binding case |
|---|---|---|
| RPM | Requests per minute | Many small classification calls |
| ITPM | Input tokens per minute | Large prompts / long context / uncached prefixes |
| OTPM | Output tokens per minute | Long generations, streaming many tokens |
| Tier | How you move up | Notes |
|---|---|---|
| Tier 1–4 (standard) | Automatic with usage + payment history | Higher tiers raise RPM/ITPM/OTPM |
| Custom / enterprise | Sales agreement | Committed throughput |
| Priority Tier | Reserved capacity for latency-sensitive traffic | Excludes Fable 5.1, Opus 5, Sonnet 5; Haiku 4.5 eligible |
Exam signal
A 429 with lots of cached input still counts cached-read tokens against ITPM at a reduced rate but writes count in full — caching lowers cost more than it lowers ITPM pressure. If a stem says “we keep hitting 429 on a 300k-token prompt at low request volume”, the binding limit is ITPM, not RPM: shrink the prompt, cache the prefix, or raise the tier.
Deprecation and migration lifecycle
Model retirement is a planned lifecycle event, not an emergency. The exam-correct posture: watch deprecation notices, keep an eval set, migrate behind that eval, pin IDs in production.
Announced → Deprecated (still callable) → Retirement floor date → Retired (404) │ │ │ │ └─ start eval-gated migration └─ published earliest date; often extended └─ appears in models.list() deprecation metadata / dashboard| Reasoning step | What to do |
|---|---|
| Notice arrives | Read the retirement floor date; it is the earliest, not a promise |
| Pick the target | Newer-or-equal model; check the capability matrix for breaking changes |
| Guard the migration | Run the existing golden set on the new model; compare per-segment, not aggregate |
| Handle thinking blocks | Migrating up (older→newer) is safe; migrating a Fable 5.1 flow down drops thinking blocks |
| Cut over | Change the pinned ID; keep the old ID available for rollback until the floor date |
Per-model migration checklists
- To Haiku 4.5 — remove any
effortparameter (unsupported → 400); if you relied onbudget_tokens, this is the only current model that keeps it, so a downgrade from adaptive-thinking models is wherebudget_tokensreappears. Re-tune prompts for the 200k/64k window; verify classification/extraction accuracy on the golden set. - To Sonnet 5 — drop any mid-conversation
role: "system"messages and task budgets (unsupported); confirmxhigheffort is not assumed (no benefit). Re-check latency budgets; it is the balanced default. - To Opus 5 — safe target for most agentic upgrades;
xhighavailable. Removebudget_tokens(400) in favour of{"type":"adaptive"}. Re-baseline cost — 2.5× Sonnet input. - To Fable 5.1 — the highest-friction migration. Remove forced
tool_choice(any/named → 400); switch toauto+instruction,strict: true, oroutput_config.format. Freezesystem/toolsfor append-only history. Confirm the org is not ZDR (30-day retention required) and does not depend on Priority Tier. Re-baseline cost at $10/$50.
More worked cost scenarios
Example 4 – caching + Batch combined
Nightly enrichment of 50,000 records on Sonnet 5. Each request: 12,000-token shared instruction/schema prefix (cacheable) + 800 unique input tokens + 300 output tokens. Run as one Batch job.
| Cost component | Calculation | Cost |
|---|---|---|
| Cached prefix reads (Batch, 50% off the 0.1× read) | 50k × 12k × ($2 × 0.1 × 0.5)/1M = 600M tok × $0.10 | $60.00 |
| One cache write (first request seeds it) | 12k × ($2 × 1.25)/1M | $0.03 |
| Unique input (Batch 50%) | 50k × 800 × ($2 × 0.5)/1M = 40M × $1.00 | $40.00 |
| Output (Batch 50%) | 50k × 300 × ($10 × 0.5)/1M = 15M × $5.00 | $75.00 |
| Total | ≈ $175.03 |
Naïve Sonnet 5 realtime, no cache: prefix 50k×12k=600M×$2=$1,200 + input 40M×$2=$80 + output 15M×$10=$150 = $1,430. Caching+Batch cuts it ≈ 88%. The prefix dominates, so caching it matters far more than the discount on the small unique portion.
Cache scope in Batch
Cache hits require the prefix to be identical and within the TTL window. In a large Batch the writes/reads interleave across the job; budget for a handful of writes, not one. The dominant saving still comes from reading the 12k prefix 50,000 times at 0.1×.
Example 5 – cascade routing math
A stream of 100,000 tickets. A Haiku 4.5 first pass answers 100% (1,500 in / 400 out each). An external validator flags 18% as low-confidence; those escalate to Opus 5 (same tokens). Compare to sending everything to Opus 5.
| Path | Input | Output | Cost |
|---|---|---|---|
| Haiku pass (all 100k) | 150M × $1 = $150 | 40M × $5 = $200 | $350 |
| Opus escalation (18k) | 27M × $5 = $135 | 7.2M × $25 = $180 | $315 |
| Cascade total | $665 | ||
| Opus-only (all 100k) | 150M × $5 = $750 | 40M × $25 = $1,000 | $1,750 |
Cascade saves ≈ 62%. The break-even escalation rate e where cascade = Opus-only solves 350 + 1750e = 1750 → e ≈ 80%. Below ~80% escalation, cascade wins; above it, the Haiku pass is pure overhead — send everything to Opus.
Exam signal
Escalation must be triggered by an external validator or a downstream check, never by the cheap model’s self-reported confidence (self-report reliance, anti-pattern 4). A stem that routes on “if Haiku says it is unsure” is the distractor; “if a validator rejects the answer” is correct.
Common misconceptions
| Misconception | Reality | Why it matters on the exam |
|---|---|---|
| “Fable 5.1 is just a bigger Opus 5, use it everywhere” | It is most capable but $10/$50, no forced tools, no ZDR, not Priority Tier | Cost-blind and constraint-blind distractor |
| “Prompt caching makes requests free after the first” | Reads are 0.1×, not 0; writes cost 1.25×/2×; TTL expires | Overstates savings; break-even ≈ 2 reads |
| “Batch is slower so it is worse” | Batch trades latency for 50% cost; ideal for offline/overnight | Latency-blind vs cost-sensitive stems |
| “Temperature 0 makes Claude deterministic and correct” | Reduces variance, not error; not a truth switch | Confuses variance with accuracy |
| “Pick the newest model to be safe” | Newest can add breaking changes (Fable 5.1) and cost 5–10× | Over-engineered / cost-blind |
| “Rate limits are just requests per minute” | RPM, ITPM and OTPM bind independently | A large-prompt 429 is ITPM, not RPM |
| “A retirement date means the model dies then” | It is the earliest floor; migrate behind an eval before it | Panic-migration distractor |
Scenario walkthrough
A fintech runs a document-classification pipeline on 2M PDFs/month. Requirements: US-federal customer (FedRAMP High), personal data (needs regional residency and, ideally, ZDR), cost-sensitive, latency non-critical (nightly), classification quality must be measured per document type before switching models.
Expert reasoning trace:
- Residency + FedRAMP High rules out the direct API for the regulated workload → Bedrock or Vertex. Pick the one matching existing cloud spend/IAM.
- ZDR desired + cheap + high-volume classification → Haiku 4.5. It is Priority-Tier eligible (irrelevant here, latency non-critical) and supports ZDR. Fable 5.1 is rejected: no ZDR, 10× the price, and classification does not need frontier reasoning (constraint-blind + over-engineered).
- Latency non-critical, 2M/month → Batch API (50% off). Realtime is rejected as needlessly expensive (cost-blind).
- Shared schema/instruction prefix → prompt caching on the stable prefix; dynamic PDF last.
- “Measured per document type” → per-segment eval on a golden set, not aggregate accuracy (aggregate-metric distractor).
- Model migration later → pin
claude-haiku-4-5, keep the eval set, migrate behind it when a successor ships.
Correct architecture: Bedrock + Haiku 4.5 + Batch + prompt caching + per-segment evals. Each tempting alternative (Fable everywhere, realtime for speed, aggregate accuracy, self-reported confidence routing) maps to a named distractor pattern.
Last updated Sep 18, 2026