Domains
D2 · Claude Models, Prompting & Context Engineering
Portfolio-level model selection and routing, prompts and guardrails as governed assets, thinking/effort control, context and token optimisation, caching architecture, prompt versioning, and Fable 5.1 harness constraints.
This domain is worth 13% – roughly 8 of 63 items. At Professional level you are not writing a single prompt; you are managing a portfolio of models and a library of governed prompts across an organisation, with cost math, versioning, rollout and breaking-change management. Items reward the option that treats prompts, models and thinking configuration as engineered, versioned assets with measurable trade-offs.
Learning objectives
By the end of this page you should be able to:
- Select and route/cascade across the model portfolio (Fable 5.1, Opus 5, Sonnet 5, Haiku 4.5) with real cost math.
- Manage breaking changes across model versions (thinking-block rules, removed parameters).
- Treat system prompts, templates and guardrails as governed, versioned assets.
- Apply prompting techniques (zero-shot, few-shot, CoT, extended/adaptive thinking, effort) at the right altitude.
- Optimise context window and tokens.
- Design a caching architecture (stable-prefix-first, modular prompts, Skills) and a versioning/rollout process.
- Respect Fable 5.1’s append-only harness constraints.
2.1 The model portfolio and routing
There is no single “best” model — there is a portfolio, and the architect’s job is to route each request to the cheapest model that meets its quality bar.
| Model | ID | Context / max out | In / out per MTok | When it is the right default |
|---|---|---|---|---|
| Fable 5.1 | claude-fable-5-1 | 1M / 128k | $10 / $50 | Frontier reasoning; thinking always on; append-only harness; 30-day retention |
| Opus 5 | claude-opus-5 | 1M / 128k | $5 / $25 | Complex agentic coding and enterprise work |
| Sonnet 5 | claude-sonnet-5 | 1M / 128k | $2 / $10 | Best speed/intelligence balance; general default |
| Haiku 4.5 | claude-haiku-4-5 | 200k / 64k | $1 / $5 | Fastest/cheapest; high-volume, latency-sensitive, narrow tasks |
Routing and cascades
┌──────────────┐ classify task difficulty / stakes request ────▶ │ router │ (rules, or a Haiku classifier) └──────┬───────┘ easy / high-vol │ hard / high-stakes ▼ ▼ ┌────────────┐ ┌────────────┐ │ Haiku 4.5 │ │ Opus 5 / │ │ Sonnet 5 │──fail─▶│ Fable 5.1 │ (escalate on └────────────┘ check └────────────┘ validation failure)- Static routing: rules by task type/segment (extraction → Haiku, synthesis → Sonnet, hardest coding → Opus/Fable).
- Cascade: try cheap first; escalate on a validation/quality check failure, not on self-reported confidence.
Exam signal
“Cost is 4× budget but quality is fine” → introduce a cascade (cheap-first, escalate on failed checks). “We need the newest capability and cost is secondary” → Opus 5 / Fable 5.1. Distractors that escalate on self-reported confidence are wrong (anti-pattern #4).
Cost math worked example
100k requests/day, ~3k input + 1k output tokens each.
| Strategy | Daily cost (approx.) | Note |
|---|---|---|
| All Opus 5 | 100k × (3k×$5 + 1k×$25)/1e6 = 100k × $0.040 = $4,000 | baseline |
| All Sonnet 5 | 100k × (3k×$2 + 1k×$10)/1e6 = 100k × $0.016 = $1,600 | 60% cheaper |
| Cascade: 80% Haiku, 20% Opus | 80k×$0.008 + 20k×$0.040 = $1,440 | quality preserved on hard 20% |
| Sonnet + cache (80% cache hit on 2k prefix) | ~$960 | cache read ≈ 0.1× input |
2.2 Managing breaking changes across versions
Model IDs from 4.6 onward are dateless but still pinned snapshots. Upgrading a model can change behaviour and can break code that relied on removed parameters.
| Change | Affected | Migration action |
|---|---|---|
budget_tokens removed (returns 400) | Fable 5.x, Opus 5, Sonnet 5, Opus 4.7–4.8 | Use thinking: {"type": "adaptive"} and effort; only Haiku 4.5 still uses budget_tokens |
| Forced tool use returns 400 on Fable 5.1 | Fable 5.1 | tool_choice:"any" / forced {"type":"tool"} → use auto + instruction, strict:true, or structured outputs |
| Thinking blocks readable only by producing model or newer | all thinking models | Never silently fall back to an older model mid-session |
| Opus 4.1 retired 2026-08-05 | pinned to 4.1 | Repin and re-run the eval/regression suite |
Always re-run the regression suite (D4) before promoting a model change. Use client.models.list() / .retrieve(id) to confirm live limits.
2.3 Prompts and guardrails as governed assets
At Professional level a system prompt is a shared organisational asset, not a string in someone’s notebook. Govern it like code.
| Property | Practice |
|---|---|
| Versioned | Stored in VCS with semantic version; changes reviewed |
| Testable | Each version runs against the golden eval set before release |
| Modular | Composed of stable blocks (role, policy, format) + volatile blocks (task) |
| Guardrailed | Safety and business rules layered, not buried in prose |
| Rolled out | Canary → percentage → full, with rollback |
Prompt-as-enforcement anti-pattern
Critical business rules (“never issue a refund over $500”) must be enforced programmatically (tool permission hooks, validation), not by a sentence in the system prompt. A model can be talked out of prose rules; a hook cannot. Options that rely on prompt wording to enforce a hard rule are wrong (anti-pattern #3).
2.4 Prompting techniques at the right altitude
| Technique | Use when | Cost/latency note |
|---|---|---|
| Zero-shot | Task is common and well-specified | Cheapest |
| Few-shot | Output shape/format must be pinned; edge cases shown | Adds input tokens (cache the examples) |
| Chain-of-thought | Reasoning must be explicit (older/non-thinking models) | More output tokens |
| Extended / adaptive thinking | Hard reasoning; thinking:{"type":"adaptive"} on current models | Thinking tokens billed as output |
Effort (low|medium|high|xhigh) | Tune reasoning depth vs cost; xhigh for hardest coding/agentic on Opus 5 / Fable 5.1 | Higher effort = more tokens/latency |
| Fast mode | Latency-sensitive use | Trades some depth for speed |
On current models, prefer adaptive thinking + effort over manual budgets. Only Haiku 4.5 still accepts budget_tokens and has **no effort parameter`.
{ "model": "claude-opus-5", "thinking": { "type": "adaptive" }, "effort": "xhigh", "messages": [{ "role": "user", "content": "Refactor this service for idempotency." }]}2.5 Context-window and token optimisation
A 1M-token window is a budget, not a target. Filling it raises cost and latency and can reduce accuracy (needle-in-haystack degradation).
| Lever | Effect |
|---|---|
| Retrieve, don’t stuff | Send only relevant chunks (RAG, D3) instead of whole corpora |
| Prompt caching | Amortise stable prefix; cache read ≈ 0.1× input |
| Context editing | Clear stale tool results from the window |
| Compaction | Server-side summarisation preserving narrative for long sessions |
| Output trimming | Ask for the schema you need; avoid verbose prose |
| Structured outputs | Fewer wasted tokens than free-form + reparse |
2.6 Caching architecture: stable-prefix-first
Prompt caching only helps if the cache prefix is stable. Order content most-stable first: system prompt → tools → long documents → few-shot examples → volatile user turn. Mark the stable boundary with cache_control.
[ system prompt ] ← stable ┐[ tool definitions ] ← stable │ cache_control: {"type":"ephemeral"} (this prefix is cached)[ reference documents]← stable │[ few-shot examples ] ← stable ┘------------------------------------- cache boundary[ user's actual question ] ← volatile (changes every request → never cache here)- Cache write ≈ 1.25× (5-min TTL) or 2× (1-hour TTL); cache read ≈ 0.1× base input.
- Minimum cacheable prefix ~1024 tokens (2048 on Haiku).
- Putting anything volatile before the stable content invalidates the cache on every call — a classic mistake.
Modular prompts + Skills: keep reusable capability blocks as Skills (SKILL.md, loaded progressively on demand) rather than pasting everything into every prompt. This keeps the cacheable prefix stable and the context lean.
2.7 Prompt versioning and rollout
Treat each prompt as name@semver. Store in VCS. Record which model version it was validated against — a prompt tuned for Opus 5 is not guaranteed to behave on Haiku 4.5. Tag the eval scores achieved.
Canary the new version on a small traffic slice with online metrics and a guardrail on regression. Ramp by percentage. Keep the previous version hot for instant rollback. Never swap a prompt org-wide without an offline eval pass first.
Because prompts are versioned and the prior version stays deployable, rollback is a config flip, not a redeploy. This is why prompts belong in a governed registry, not inline in application code.
2.8 Fable 5.1 append-only harness constraints
Fable 5.1 has thinking always on, and its thinking blocks are readable only by the producing model (or a newer one). Editing, reordering or removing earlier turns invalidates later thinking blocks. Therefore harnesses must be append-only.
| Rule | Consequence if violated |
|---|---|
Freeze system and tools after the session starts | Editing them invalidates downstream thinking |
Put mid-session changes in a role: "system" message (append, don’t edit) | Rewriting history breaks the loop |
| Trim server-side via context editing / compaction, not by deleting turns client-side | Client-side deletion invalidates thinking blocks |
Never force tool use (tool_choice:"any" / forced tool) — returns 400 | Request fails; use auto + instruction / strict / structured outputs |
| Never silently fall back to an older model | Older model drops the thinking blocks |
Sonnet 5 differs
Sonnet 5 does not support mid-conversation system messages and has no task budgets. Do not assume Fable’s append-a-system-message trick works identically on Sonnet 5 — the exam tests these per-model differences.
2.9 Prompt-caching cost arithmetic
Caching only pays when you can quantify it. Read ≈ 0.1× base input; write ≈ 1.25× (5-min TTL) or 2× (1-hour TTL); minimum cacheable prefix ~1024 tokens (2048 on Haiku).
Worked example. Sonnet 5, 4,000-token stable prefix ($2/MTok input), 500-token volatile turn, 10,000 requests/day, 90% cache-hit rate after warm-up.
Uncached input cost/req = 4,500 × $2 / 1e6 = $0.0090Cached (hit) input cost = (4,000 × 0.1 + 500) × $2/1e6 = (400 + 500)×$2/1e6 = $0.0018Cache write (miss, 1.25×) = (4,000 × 1.25 + 500) × $2/1e6 ≈ $0.0110 (paid on ~10% of calls)
Daily uncached = 10,000 × $0.0090 = $90.00Daily cached = 0.9×10,000×$0.0018 + 0.1×10,000×$0.0110 = $16.20 + $11.00 = $27.20Saving ≈ 70% of input cost.Exam signal
Caching helps in proportion to prefix size × hit rate. A tiny prefix or a low hit rate (because a volatile token sits in the prefix) makes caching worthless. If a stem says “hit rate is near zero”, look for a volatile element before the cache boundary — not “disable caching”.
| Symptom | Cause | Fix |
|---|---|---|
| Hit rate near zero | Volatile content before the boundary; per-request timestamp/user-id in prefix | Move volatile content after the cache_control boundary |
| Write cost dominates | Prefix rarely reused within the TTL | Use the 1-hour TTL, or don’t cache low-reuse prefixes |
| No effect on Haiku | Prefix under the 2048-token minimum | Consolidate stable context or accept no caching |
2.10 Structured outputs and schema enforcement
At Professional level, “parse the prose” is never the answer. Use output_config.format with a JSON schema and strict: true so the model’s output conforms by construction, and reserve validation-retry for the rare miss.
{ "model": "claude-sonnet-5", "messages": [{ "role": "user", "content": "Extract the invoice fields." }], "output_config": { "format": { "type": "json_schema", "schema": { "type": "object", "properties": { "invoice_id": { "type": "string" }, "total": { "type": "number" }, "currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] } }, "required": ["invoice_id", "total", "currency"], "additionalProperties": false }, "strict": true } }}| Approach | Reliability | When |
|---|---|---|
| Free-text + regex/parse | Brittle | Never for structured data |
| Prompt “return JSON” only | Better, still fallible | Legacy/unsupported paths |
strict JSON schema (structured outputs) | Conforms by construction | Default for machine-consumed output |
Tool schema with strict: true | Enforced tool arguments | When a tool needs typed args |
Exam signal
On Fable 5.1 you cannot force tool use (tool_choice:"any" → 400). To guarantee a shape, use structured outputs / strict schema or auto + instruction — the exam pairs the “guaranteed JSON” need with the “no forced tools on Fable” constraint.
2.11 Context editing vs compaction
Long-running sessions overflow the window. Two server-side tools manage it, and they are not interchangeable.
| Technique | What it does | Use when | Risk if misused |
|---|---|---|---|
| Context editing | Removes/clears stale tool results and blocks from the window | Tool outputs are large and no longer needed | Editing earlier turns invalidates Fable 5.1 thinking blocks — edit tool results, not reasoning |
| Compaction | Server-side summarisation preserving the narrative | Very long sessions where history must be retained in gist | Over-compaction loses detail needed later |
| Memory tool | Persist durable facts outside the window | Facts must survive across sessions | Storing secrets/PII inappropriately |
Prefer trimming server-side (context editing / compaction) over client-side deletion of turns, which breaks append-only harness invariants (2.8). The PreCompact hook lets you snapshot state before compaction runs.
2.12 Scenario walkthrough: taming a runaway prompt-and-model bill
Scenario. A document-analysis product runs every request on Opus 5 with effort: xhigh, a 9k-token system prompt duplicated per call, and forced tool use. Monthly spend is 5× budget; the team also just failed a Fable 5.1 pilot with 400 errors. Quality is acceptable; latency is not the complaint — cost is.
Expert reasoning trace.
-
Right-size the model. Quality is already acceptable on Opus 5, so most traffic can run on Sonnet 5 with a cascade escalating to Opus 5 only on a validation-check failure. That alone cuts the per-request rate ~60%.
-
Fix the effort.
xhigheverywhere is wasteful; drop tohigh/mediumand re-run the regression suite per segment to confirm no drop. -
Cache the prefix. The 9k-token system prompt is stable → mark a
cache_controlboundary; move the per-request document after it. Reads at 0.1× turn the prefix nearly free at a high hit rate. -
Modularise with Skills. The duplicated capability text belongs in Skills loaded on demand, keeping the cached prefix lean and stable.
-
Explain the Fable 400s. Forced tool use is unsupported on Fable 5.1; switch to
auto+ instruction or structured outputs. But note Fable’s $10/$50 pricing makes it the wrong cost choice here anyway. -
Re-validate and roll out. Offline regression per segment → canary → ramp, prior version hot for rollback.
Why the tempting alternatives are wrong: “move everything to Haiku” risks the quality bar; “escalate on self-reported confidence” is anti-pattern #4; “just buy a bigger budget” ignores the arithmetic; “keep forcing tools and retry” cannot fix a 400.
2.13 Common misconceptions
| Misconception | Reality | Why it matters on the exam |
|---|---|---|
“budget_tokens is the standard way to control thinking.” | Removed on current models (400); only Haiku 4.5 still uses it. Use adaptive thinking + effort. | A 400-after-upgrade stem tests exactly this. |
| “A strong system-prompt sentence enforces a business rule.” | Prompts are guidance; hard rules need hooks/validation. | Prompt-as-enforcement is a recurring wrong answer. |
| “Higher effort always means better answers.” | Beyond the task’s need it just adds tokens/latency/cost. | xhigh-everywhere distractors overspend. |
| “Caching automatically saves money once enabled.” | Only if the prefix is stable and reused above the minimum size. | Volatile-prefix stems make caching worthless. |
| “Forcing tool use works on every model.” | Fable 5.1 returns 400 on forced tool use. | Guaranteed-shape stems pair with structured outputs. |
| “A prompt tuned on Opus behaves the same on Haiku.” | Behaviour differs per model; re-validate per model. | Cross-model reuse without re-eval is the trap. |
| “You can trim a long session by deleting old turns client-side.” | On thinking models that invalidates later thinking blocks; trim server-side. | Append-only harness rules are tested per model. |
Exam traps in this domain
| Trap | Why it is wrong |
|---|---|
| Using the most expensive model for every request | Ignores routing/cascades; blows the cost budget |
| Escalating in a cascade on self-reported confidence | Self-report is unreliable (anti-pattern #4); escalate on validation failure |
| Enforcing a hard business rule via the system prompt | Prompt-as-enforcement (anti-pattern #3); use hooks/validation |
Setting budget_tokens on Opus 5 / Sonnet 5 / Fable 5.1 | Removed; returns 400 — use adaptive thinking + effort |
| Forcing tool use on Fable 5.1 | Returns 400; use auto+instruction, strict, or structured outputs |
| Editing earlier turns in a Fable 5.1 session | Invalidates later thinking blocks; harness must be append-only |
| Silently falling back to an older model mid-session | Drops thinking blocks; corrupts the session |
| Putting the volatile user turn before the cached prefix | Invalidates the cache every call |
| Filling the 1M window “because it’s available” | Raises cost/latency; can reduce accuracy; retrieve instead |
| Assuming a prompt tuned on one model behaves identically on another | Must re-validate per model version |
Parsing free-text prose for structured data instead of using a strict schema | Brittle; structured outputs conform by construction |
| Deleting old turns client-side to trim a thinking-model session | Invalidates later thinking blocks; use context editing/compaction |
| Assuming caching saves money regardless of prefix size or hit rate | Saving ∝ prefix size × hit rate; a volatile prefix yields ~0 |
| Storing secrets/PII in the memory tool or persisted context | Exfiltration/compliance risk; keep secrets in secret managers |
Raising effort to xhigh to “improve quality” without evidence | Adds tokens/latency/cost past the task’s need |
Practice questions
Q1 · A pipeline routes everything to Opus 5. Cost is 4× budget; quality is acceptable. Which change best cuts cost while preserving quality on hard cases? (Select one)
A. Move all traffic to Haiku 4.5. B. Build a cascade: Haiku 4.5 / Sonnet 5 first, escalate to Opus 5 only when an output validation check fails, and cache the stable prefix. C. Escalate to Opus 5 whenever the model reports low confidence in its own answer. D. Increase effort to xhigh everywhere.
Answer: B. Cheap-first with escalation on validation failure preserves quality on hard cases while most traffic runs cheaply; caching amortises the stable prefix. Blanket Haiku (A) sacrifices quality. Self-reported confidence (C) is an anti-pattern. Raising effort everywhere (D) increases cost.
Q2 · A team upgrades from Opus 4.6 to Opus 5 and their requests now return 400 errors. The requests set `budget_tokens` for thinking. What is the fix? (Select one)
A. Add more retries.
B. Replace budget_tokens with thinking: {\"type\": \"adaptive\"} and control depth via effort, since budget_tokens is removed on Opus 5.
C. Downgrade permanently to Haiku 4.5.
D. Remove thinking entirely.
Answer: B. budget_tokens is removed on Opus 5 (and Sonnet 5 / Fable 5.x) and returns 400; adaptive thinking plus effort is the supported replacement. Retries (A) won’t fix a 400. Haiku (C) is the only model still using budget_tokens but is not an equivalent for Opus workloads. Removing thinking (D) discards needed reasoning.
Q3 · A refund agent must never issue refunds above $500. Where should this rule live? (Select one)
A. As a firmly worded sentence in the system prompt. B. As a programmatic tool-permission hook / validation that rejects any refund over $500 before execution. C. As a few-shot example showing a refused large refund. D. In the model’s thinking budget.
Answer: B. Hard business rules require programmatic enforcement — a hook or validation the model cannot talk its way past. Prompt wording (A) and few-shot examples (C) are prompt-as-enforcement anti-patterns. Thinking budget (D) is unrelated.
Q4 · Prompt caching is enabled but hit rate is near zero. The prompt places the user's question first, then the system prompt and reference documents. Why, and what fixes it? (Select one)
A. Caching is broken; disable it.
B. The volatile user turn sits before the stable content, so the cached prefix changes every call — reorder to stable-first (system → tools → docs) then the user turn, and mark the stable boundary with cache_control.
C. The documents are too short.
D. Haiku doesn’t support caching.
Answer: B. Caching keys on a stable prefix; putting the changing user turn first invalidates it every call. Stable-prefix-first with a cache_control boundary fixes it. Caching is not broken (A); document length (C) matters only for the ~1024-token minimum; Haiku does support caching (D) with a 2048-token minimum.
Q5 · Which are valid reasons NOT to stuff a whole 500k-token corpus into the 1M window every request? (Select two)
A. Higher token cost and latency per call. B. Possible accuracy degradation locating the relevant needle. C. The window physically cannot hold it. D. Structured outputs are disabled above 200k tokens. E. Caching is prohibited on large inputs.
Answer: A and B. Stuffing raises cost and latency and can hurt retrieval accuracy within a huge context; retrieval (RAG) sends only relevant chunks. It fits the window (C is false at 500k of 1M), structured outputs are not size-gated that way (D), and caching is allowed on large inputs (E).
Q6 · In an active Fable 5.1 agentic session, the team wants to change the system instructions mid-run. What is the correct approach? (Select one)
A. Edit the original system field in place.
B. Append the change as a new role: \"system\" message, leaving earlier turns untouched, because the harness must be append-only.
C. Delete the earliest turns to make room.
D. Reorder messages to put the new instruction first.
Answer: B. Fable 5.1 thinking blocks are invalidated by editing/reordering/removing earlier turns, so mid-session changes are appended as a new system message. Editing in place (A), deleting turns (C) and reordering (D) all invalidate downstream thinking blocks.
Q7 · A team wants to force Fable 5.1 to always return a tool call using tool_choice set to any. It returns 400. What should they do? (Select one)
A. Retry until it works.
B. Use tool_choice: \"auto\" with an instruction to use the tool, set strict: true on the tool schema, or use structured outputs — because forced tool use is unsupported on Fable 5.1.
C. Switch to Haiku 4.5 permanently.
D. Remove all tools.
Answer: B. Forced tool use (any / forced tool) returns 400 on Fable 5.1; the supported paths are auto + instruction, strict schemas, or structured outputs. Retrying (A) won’t fix a 400. Switching model (C) or removing tools (D) abandons the requirement.
Q8 · How should a new system-prompt version be rolled out to production? (Select one)
A. Swap it org-wide immediately to move fast. B. Validate offline against the golden set, canary on a small traffic slice with regression guardrails, ramp by percentage, and keep the prior version hot for instant rollback. C. Let each engineer edit the inline prompt in their own service. D. Ship it and monitor customer complaints.
Answer: B. Prompts are governed, versioned assets: offline eval → canary → ramp → keep prior version for rollback. Org-wide swaps (A) and complaint-driven monitoring (D) skip validation. Per-engineer inline edits (C) destroy governance and reproducibility.
Q9 · A high-volume extraction subtask feeds a slower synthesis step. Which portfolio assignment is BEST? (Select one)
A. Opus 5 for both steps. B. Haiku 4.5 for the high-volume extraction; Sonnet 5 or Opus 5 for the synthesis — matching model cost to each step’s difficulty. C. Fable 5.1 for both, for maximum quality. D. Haiku 4.5 for both, for maximum savings.
Answer: B. Portfolio routing assigns the cheapest adequate model per step: Haiku for narrow high-volume extraction, a stronger model for harder synthesis. Opus/Fable for both (A, C) overpays; Haiku for both (D) risks the synthesis quality.
Q10 · Reusable capability blocks are being pasted into every prompt, bloating context and breaking the cache prefix. What is the better pattern? (Select one)
A. Package them as Skills (SKILL.md) loaded progressively on demand, keeping the stable cacheable prefix lean.
B. Duplicate them into each service’s prompt.
C. Move them into the volatile user turn.
D. Increase the context window.
Answer: A. Skills load capability progressively on demand, keeping context lean and the cache prefix stable. Duplication (B) is what caused the bloat; moving them to the volatile turn (C) worsens caching; a bigger window (D) doesn’t address cost or cache stability.
Q11 · Which statement about current-model thinking configuration is correct? (Select one)
A. All current models require budget_tokens.
B. Current models use thinking: {\"type\":\"adaptive\"} with effort levels low/medium/high/xhigh; only Haiku 4.5 still uses budget_tokens and has no effort parameter.
C. Effort only exists on Haiku 4.5.
D. xhigh effort is the default on all models.
Answer: B. Adaptive thinking plus effort is standard on current models; Haiku 4.5 is the exception still using budget_tokens and lacking effort. budget_tokens is not universal (A); effort is not Haiku-only (C); high (not xhigh) is the default (D).
Q12 · A prompt validated on Opus 5 is reused verbatim on Haiku 4.5 and quality drops. What is the correct lesson? (Select one)
A. Haiku 4.5 is defective. B. Prompts are validated per model version; a prompt tuned for one model must be re-tested (and often adjusted) against the golden set on any other model before use. C. Always use Opus 5. D. Quality drops are unavoidable and should be ignored.
Answer: B. Model behaviour differs across the portfolio, so prompt versions carry the model they were validated against and must be re-evaluated when reused elsewhere. Haiku is not defective (A); mandating Opus (C) ignores cost; ignoring regressions (D) is negligent.
Q13 · A 4,000-token stable prefix on Sonnet 5 is reused with a 90% cache-hit rate at 10,000 req/day. Roughly what does caching save on input cost? (Select one)
A. Nothing; caching never helps large prefixes. B. Around 70%, because a cache read is ~0.1× input so the 4k prefix becomes ~400 effective tokens on hits. C. Exactly 50%, the Batch discount. D. 100%; cached requests are free.
Answer: B. Read ≈ 0.1× input turns 4,000 prefix tokens into ~400 on the 90% of hits; the daily input cost drops from ~$90 to ~$27, about 70%. Caching does help large stable prefixes (A); 50% is the Batch discount, not caching (C); cached reads are cheap, not free (D).
Q14 · A pipeline must return a strictly typed JSON object for downstream systems, and it runs on Fable 5.1 where forcing tool use returns 400. What is the BEST approach? (Select one)
A. Force a tool call with tool_choice: 'any' and retry on 400.
B. Use structured outputs with a strict JSON schema (or auto + instruction), which guarantees the shape without forcing tool use.
C. Ask for JSON in the prompt and regex-parse the prose.
D. Switch to free-text and reparse.
Answer: B. Structured outputs with a strict schema conform by construction and don’t require forced tool use, which Fable 5.1 rejects with 400. Forcing tools (A) fails; prompt-only JSON with regex (C) and free-text reparse (D) are brittle prose-parsing.
Q15 · A long agentic session on a thinking model overflows the window because tool results are huge. Which technique is correct, and what must be avoided? (Select one)
A. Delete the earliest user/assistant turns client-side. B. Use context editing to clear stale tool results server-side (and compaction for narrative), avoiding client-side edits to earlier reasoning that invalidate thinking blocks. C. Lower the temperature to shrink the context. D. Force the model to summarise itself in the same turn.
Answer: B. Context editing removes stale tool results server-side; compaction summarises narrative — both avoid invalidating thinking blocks. Deleting turns client-side (A) breaks the append-only invariant; temperature (C) doesn’t affect context size; in-turn self-summary (D) doesn’t reclaim the window.
Q16 · A team runs every request on Opus 5 at `effort: xhigh` with acceptable quality and 5× budget; cost, not latency, is the complaint. Which TWO changes best cut cost while preserving quality? (Select two)
A. Cascade most traffic to Sonnet 5, escalating to Opus 5 only on a validation-check failure. B. Lower effort to an adequate level and re-run the per-segment regression suite. C. Escalate on the model’s self-reported confidence. D. Move all traffic to Fable 5.1 for quality. E. Remove the eval suite to cut compute.
Answer: A and B. A cheap-first cascade and right-sized effort (re-validated per segment) cut cost while preserving quality on hard cases. Self-report (C) is anti-pattern #4; Fable 5.1 (D) is the most expensive model; removing evals (E) removes the quality guard.
Q17 · Caching is enabled but the hit rate is ~3%. Investigation shows a per-request `request_id` string is prepended to the system prompt. What is the fix? (Select one)
A. Disable caching; it doesn’t work here.
B. Remove the volatile request_id from the prefix (log it separately) so the prefix is byte-stable, restoring cache hits.
C. Shorten the documents.
D. Increase the context window.
Answer: B. A per-request token in the prefix changes it every call, so nothing caches; moving it out restores a stable prefix. Caching isn’t broken (A); document length (C) only affects the minimum; a bigger window (D) is unrelated.
Q18 · A durable fact (a customer's contract tier) must persist across separate sessions without re-sending it in every prompt. Which mechanism fits, and what constraint applies? (Select one)
A. Paste the fact into every system prompt.
B. Use the memory tool to persist the fact across sessions, but never store secrets/PII there inappropriately and keep it out of model-visible logs.
C. Store it in CLAUDE.local.md.
D. Increase retention to keep it in provider logs.
Answer: B. The memory tool persists durable facts across sessions; the constraint is not to store secrets/PII inappropriately. Pasting per prompt (A) bloats context; CLAUDE.local.md (C) is a Claude Code dev file, not a runtime store; relying on provider retention (D) is not a memory mechanism and raises compliance risk.
Key takeaways
- Manage a portfolio: route/cascade to the cheapest model that clears the quality bar; escalate on validation failure, not self-reported confidence.
- Do the cost math — cascades and caching routinely cut spend 60–75% with quality preserved on hard cases.
- Manage breaking changes:
budget_tokensremoved (400) except on Haiku 4.5; forced tool use fails on Fable 5.1; re-run regressions before promoting a model. - Treat prompts and guardrails as versioned, governed assets; enforce hard rules programmatically, never via prompt prose.
- Use adaptive thinking + effort; reserve
xhighfor the hardest Opus/Fable work. - Cache stable-prefix-first; keep reusable blocks as Skills; never place volatile content before the cached prefix.
- Fable 5.1 harnesses are append-only: freeze system/tools, append system messages, trim server-side, never force tools or silently downgrade.
- Caching saving ∝ prefix size × hit rate; a volatile token in the prefix (timestamps, request IDs) drops the hit rate to ~0 — remove it, don’t disable caching.
- Guarantee output shape with structured outputs /
strictschemas, especially on Fable 5.1 where forced tool use returns 400 — never parse prose. - Manage long sessions with context editing (clear stale tool results) and compaction (summarise narrative) server-side; the memory tool persists durable facts (no secrets/PII).
Last updated Sep 18, 2026