AI Cert Prep
Type to search documentation.

Appendix · Claude

Model Lineup & Pricing

Current Claude models, IDs, limits, prices, breaking changes and worked cost calculations as of September 2026.

Current lineup (September 2026)

ModelIDContextMax outputInput / Output per MTokCache readPositioning
Claude Fable 5.1claude-fable-5-11M128k$10 / $50$0.25Most capable; hardest reasoning and agentic work
Claude Opus 5claude-opus-51M128k$5 / $25$0.50Default for complex agentic coding and enterprise
Claude Sonnet 5claude-sonnet-51M128k$2 / $10$0.20Best speed / intelligence balance
Claude Haiku 4.5claude-haiku-4-5200k64k$1 / $5$0.10Fastest and cheapest; classification, routing, extraction

Cache writes cost ≈ 1.25× base input (5-minute TTL) or 2× (1-hour TTL). Batch API: 50% off input and output.

Legacy, still available: Fable 5, Opus 4.8, 4.7, 4.6, 4.5, Sonnet 4.6, 4.5. Retired: Opus 4.1 (2026-08-05).

Published retirement floors: Fable 5.1 not before 2027-09-01 · Opus 5 not before 2027-07-24 · Sonnet 5 not before 2027-06-30 · Haiku 4.5 not before 2026-10-15.

Verify live

Model IDs from the 4.6 generation onward are dateless but still pinned snapshots. Query client.models.list() and client.models.retrieve(id) for live max_input_tokens, max_tokens and capabilities rather than trusting any static table – including this one.

Capability matrix

CapabilityFable 5.1Opus 5Sonnet 5Haiku 4.5
Adaptive thinking {"type":"adaptive"}✓ (always on)✓✓–
budget_tokens thinking400 error400 error400 error✓ (only model)
effort low/medium/high/xhigh✓✓✓ (no xhigh benefit)–
tool_choice: "any" / forced tool400 error✓✓✓
Structured outputs output_config.format✓✓✓✓
strict: true tool schemas✓✓✓✓
Mid-conversation role: "system" messages✓✓––
Task budgets✓✓––
Thinking blocks portable to other models– (bound)✓✓n/a
Zero Data Retention– (30-day required)✓✓✓
Priority Tier–––✓
Prompt caching✓✓✓✓ (2048-token minimum)
Batch API✓✓✓✓

Fable 5.1 breaking changes (exam favourites)

  1. No forced tool use. tool_choice: "any" and {"type": "tool", "name": …} return 400. Alternatives: auto plus an explicit instruction; strict: true on the tool schema; output_config.format structured output.
  2. Thinking-block binding. Thinking blocks are readable only by the model that produced them or a newer one. A fallback or router switch to an older model silently drops them (unbilled) and the older model re-plans from scratch.
  3. Append-only history. Editing, reordering or removing earlier turns invalidates every later thinking block. Freeze system and tools, deliver mid-session instruction changes as role: "system" messages, and trim context server-side (context editing / compaction) rather than on the client.
  4. Retention. Requires 30-day data retention; not available to ZDR organisations unless authorised. Excluded from Priority Tier (as are Opus 5 and Sonnet 5).

Selecting a model – decision table

Signal in the scenarioChooseWhy
“Classify”, “route”, “extract simple fields”, “millions of items”, “sub-second”Haiku 4.5Cheapest, fastest; quality sufficient for narrow tasks
“General assistant”, “customer-facing chat”, “balanced cost and quality”Sonnet 5Best trade-off; 1M context
“Complex agentic coding”, “multi-step reasoning”, “enterprise default”Opus 5Default for hard agentic work
“Hardest research”, “frontier reasoning”, budget is secondaryFable 5.1Most capable; accept breaking-change constraints
“Cost matters, latency does not”Any tier + Batch API50% discount
Same long prefix reusedAny tier + prompt caching~90% saving on cached reads
Mixed difficulty streamCascade: Haiku → Sonnet → OpusCheap model handles the bulk; escalate on low confidence from a validator, not self-report

Worked cost calculations

Example 1 – 10,000 documents, 3,000 input tokens each, 500 output tokens each

ApproachInput costOutput costTotal
Haiku 4.5 realtime30M × $1 = $305M × $5 = $25$55
Sonnet 5 realtime30M × $2 = $605M × $10 = $50$110
Opus 5 realtime30M × $5 = $1505M × $25 = $125$275
Sonnet 5 Batch (50%)$30$25$55
Opus 5 Batch (50%)$75$62.50$137.50

Example 2 – prompt caching on a 20,000-token system prompt + tools, 1,000 requests/day on Sonnet 5

Without cachingWith caching (5-min TTL, ~1 write per 5 min ≈ 288 writes/day)
Prefix tokens billed20M at $2 = $40/dayWrites: 288 × 20k × $2.50 = $14.40; Reads: 712 × 20k × $0.20 = $2.85 → ≈ $17.25/day
Saving–≈ 57% on the prefix; approaches 90% as request rate rises

Rule: caching pays back after roughly two reads of the same prefix within the TTL.

Example 3 – choosing effort

TaskEffortRationale
Subagent that renames files per a speclowMechanical
Summarise a 50-page contractmediumModerate reasoning, cost-sensitive
Default production reasoninghighBaseline
Multi-hour refactor across 40 files (Opus 5 / Fable 5.1)xhighHardest agentic work

Cloud availability

PlatformNotes
Claude API (direct)Full feature surface first; simplest
Amazon BedrockAWS IAM, data residency, FedRAMP High, existing AWS spend
Google Vertex AIGCP IAM, residency, FedRAMP High
Microsoft FoundryAzure ecosystem

Feature availability can lag on cloud platforms; check per-feature docs before committing an architecture.

Feature availability by cloud platform

The exam tests the reflex that the direct Claude API gets new features first, and that a governance or residency requirement can force a platform that lags on a feature you depend on. Treat this as a snapshot to reason with, not a live SLA.

FeatureClaude APIBedrockVertex AIFoundry
Newest model on launch day✓ firstUsually days–weeks laterUsually days–weeks laterVaries
Prompt caching✓✓✓Check
Message Batches✓✓ (batch inference)✓ (batch prediction)Check
MCP connector (server-side)✓LagsLagsLags
Files API / citations✓PartialPartialCheck
Structured outputs output_config.format✓✓✓Check
Context editing / compaction✓LagsLagsLags
FedRAMP High authorisation–✓✓–
Data residency controlsLimited✓ (region)✓ (region)✓ (region)
IAM / enterprise identityAPI keysAWS IAMGCP IAMEntra ID

Exam signal

“We are on AWS with FedRAMP High and IAM already” → Bedrock, even if it means waiting for a feature. “We need the MCP connector and context editing today” → direct Claude API. A stem that pairs a residency requirement with a bleeding-edge feature is testing whether you notice the platform lag.

Rate-limit tiers and dimensions

Rate limits apply on three dimensions simultaneously; you hit whichever binds first.

DimensionMeaningTypical binding case
RPMRequests per minuteMany small classification calls
ITPMInput tokens per minuteLarge prompts / long context / uncached prefixes
OTPMOutput tokens per minuteLong generations, streaming many tokens
TierHow you move upNotes
Tier 1–4 (standard)Automatic with usage + payment historyHigher tiers raise RPM/ITPM/OTPM
Custom / enterpriseSales agreementCommitted throughput
Priority TierReserved capacity for latency-sensitive trafficExcludes Fable 5.1, Opus 5, Sonnet 5; Haiku 4.5 eligible

Exam signal

A 429 with lots of cached input still counts cached-read tokens against ITPM at a reduced rate but writes count in full — caching lowers cost more than it lowers ITPM pressure. If a stem says “we keep hitting 429 on a 300k-token prompt at low request volume”, the binding limit is ITPM, not RPM: shrink the prompt, cache the prefix, or raise the tier.

Deprecation and migration lifecycle

Model retirement is a planned lifecycle event, not an emergency. The exam-correct posture: watch deprecation notices, keep an eval set, migrate behind that eval, pin IDs in production.

text
Announced → Deprecated (still callable) → Retirement floor date → Retired (404)
│ │ │
│ └─ start eval-gated migration └─ published earliest date; often extended
└─ appears in models.list() deprecation metadata / dashboard
Reasoning stepWhat to do
Notice arrivesRead the retirement floor date; it is the earliest, not a promise
Pick the targetNewer-or-equal model; check the capability matrix for breaking changes
Guard the migrationRun the existing golden set on the new model; compare per-segment, not aggregate
Handle thinking blocksMigrating up (older→newer) is safe; migrating a Fable 5.1 flow down drops thinking blocks
Cut overChange the pinned ID; keep the old ID available for rollback until the floor date

Per-model migration checklists

  1. To Haiku 4.5 — remove any effort parameter (unsupported → 400); if you relied on budget_tokens, this is the only current model that keeps it, so a downgrade from adaptive-thinking models is where budget_tokens reappears. Re-tune prompts for the 200k/64k window; verify classification/extraction accuracy on the golden set.
  2. To Sonnet 5 — drop any mid-conversation role: "system" messages and task budgets (unsupported); confirm xhigh effort is not assumed (no benefit). Re-check latency budgets; it is the balanced default.
  3. To Opus 5 — safe target for most agentic upgrades; xhigh available. Remove budget_tokens (400) in favour of {"type":"adaptive"}. Re-baseline cost — 2.5× Sonnet input.
  4. To Fable 5.1 — the highest-friction migration. Remove forced tool_choice (any/named → 400); switch to auto+instruction, strict: true, or output_config.format. Freeze system/tools for append-only history. Confirm the org is not ZDR (30-day retention required) and does not depend on Priority Tier. Re-baseline cost at $10/$50.

More worked cost scenarios

Example 4 – caching + Batch combined

Nightly enrichment of 50,000 records on Sonnet 5. Each request: 12,000-token shared instruction/schema prefix (cacheable) + 800 unique input tokens + 300 output tokens. Run as one Batch job.

Cost componentCalculationCost
Cached prefix reads (Batch, 50% off the 0.1× read)50k × 12k × ($2 × 0.1 × 0.5)/1M = 600M tok × $0.10$60.00
One cache write (first request seeds it)12k × ($2 × 1.25)/1M$0.03
Unique input (Batch 50%)50k × 800 × ($2 × 0.5)/1M = 40M × $1.00$40.00
Output (Batch 50%)50k × 300 × ($10 × 0.5)/1M = 15M × $5.00$75.00
Total≈ $175.03

Naïve Sonnet 5 realtime, no cache: prefix 50k×12k=600M×$2=$1,200 + input 40M×$2=$80 + output 15M×$10=$150 = $1,430. Caching+Batch cuts it ≈ 88%. The prefix dominates, so caching it matters far more than the discount on the small unique portion.

Cache scope in Batch

Cache hits require the prefix to be identical and within the TTL window. In a large Batch the writes/reads interleave across the job; budget for a handful of writes, not one. The dominant saving still comes from reading the 12k prefix 50,000 times at 0.1×.

Example 5 – cascade routing math

A stream of 100,000 tickets. A Haiku 4.5 first pass answers 100% (1,500 in / 400 out each). An external validator flags 18% as low-confidence; those escalate to Opus 5 (same tokens). Compare to sending everything to Opus 5.

PathInputOutputCost
Haiku pass (all 100k)150M × $1 = $15040M × $5 = $200$350
Opus escalation (18k)27M × $5 = $1357.2M × $25 = $180$315
Cascade total$665
Opus-only (all 100k)150M × $5 = $75040M × $25 = $1,000$1,750

Cascade saves ≈ 62%. The break-even escalation rate e where cascade = Opus-only solves 350 + 1750e = 1750 → e ≈ 80%. Below ~80% escalation, cascade wins; above it, the Haiku pass is pure overhead — send everything to Opus.

Exam signal

Escalation must be triggered by an external validator or a downstream check, never by the cheap model’s self-reported confidence (self-report reliance, anti-pattern 4). A stem that routes on “if Haiku says it is unsure” is the distractor; “if a validator rejects the answer” is correct.

Common misconceptions

MisconceptionRealityWhy it matters on the exam
“Fable 5.1 is just a bigger Opus 5, use it everywhere”It is most capable but $10/$50, no forced tools, no ZDR, not Priority TierCost-blind and constraint-blind distractor
“Prompt caching makes requests free after the first”Reads are 0.1×, not 0; writes cost 1.25×/2×; TTL expiresOverstates savings; break-even ≈ 2 reads
“Batch is slower so it is worse”Batch trades latency for 50% cost; ideal for offline/overnightLatency-blind vs cost-sensitive stems
“Temperature 0 makes Claude deterministic and correct”Reduces variance, not error; not a truth switchConfuses variance with accuracy
“Pick the newest model to be safe”Newest can add breaking changes (Fable 5.1) and cost 5–10×Over-engineered / cost-blind
“Rate limits are just requests per minute”RPM, ITPM and OTPM bind independentlyA large-prompt 429 is ITPM, not RPM
“A retirement date means the model dies then”It is the earliest floor; migrate behind an eval before itPanic-migration distractor

Scenario walkthrough

A fintech runs a document-classification pipeline on 2M PDFs/month. Requirements: US-federal customer (FedRAMP High), personal data (needs regional residency and, ideally, ZDR), cost-sensitive, latency non-critical (nightly), classification quality must be measured per document type before switching models.

Expert reasoning trace:

  1. Residency + FedRAMP High rules out the direct API for the regulated workload → Bedrock or Vertex. Pick the one matching existing cloud spend/IAM.
  2. ZDR desired + cheap + high-volume classification → Haiku 4.5. It is Priority-Tier eligible (irrelevant here, latency non-critical) and supports ZDR. Fable 5.1 is rejected: no ZDR, 10× the price, and classification does not need frontier reasoning (constraint-blind + over-engineered).
  3. Latency non-critical, 2M/month → Batch API (50% off). Realtime is rejected as needlessly expensive (cost-blind).
  4. Shared schema/instruction prefix → prompt caching on the stable prefix; dynamic PDF last.
  5. “Measured per document type” → per-segment eval on a golden set, not aggregate accuracy (aggregate-metric distractor).
  6. Model migration later → pin claude-haiku-4-5, keep the eval set, migrate behind it when a successor ships.

Correct architecture: Bedrock + Haiku 4.5 + Batch + prompt caching + per-segment evals. Each tempting alternative (Fable everywhere, realtime for speed, aggregate accuracy, self-reported confidence routing) maps to a named distractor pattern.

Last updated Sep 18, 2026