AI Cert Prep
Type to search documentation.

Domains

D2 · Claude Models, Prompting & Context Engineering

Portfolio-level model selection and routing, prompts and guardrails as governed assets, thinking/effort control, context and token optimisation, caching architecture, prompt versioning, and Fable 5.1 harness constraints.

This domain is worth 13% – roughly 8 of 63 items. At Professional level you are not writing a single prompt; you are managing a portfolio of models and a library of governed prompts across an organisation, with cost math, versioning, rollout and breaking-change management. Items reward the option that treats prompts, models and thinking configuration as engineered, versioned assets with measurable trade-offs.

Learning objectives

By the end of this page you should be able to:

  1. Select and route/cascade across the model portfolio (Fable 5.1, Opus 5, Sonnet 5, Haiku 4.5) with real cost math.
  2. Manage breaking changes across model versions (thinking-block rules, removed parameters).
  3. Treat system prompts, templates and guardrails as governed, versioned assets.
  4. Apply prompting techniques (zero-shot, few-shot, CoT, extended/adaptive thinking, effort) at the right altitude.
  5. Optimise context window and tokens.
  6. Design a caching architecture (stable-prefix-first, modular prompts, Skills) and a versioning/rollout process.
  7. Respect Fable 5.1’s append-only harness constraints.

2.1 The model portfolio and routing

There is no single “best” model — there is a portfolio, and the architect’s job is to route each request to the cheapest model that meets its quality bar.

ModelIDContext / max outIn / out per MTokWhen it is the right default
Fable 5.1claude-fable-5-11M / 128k$10 / $50Frontier reasoning; thinking always on; append-only harness; 30-day retention
Opus 5claude-opus-51M / 128k$5 / $25Complex agentic coding and enterprise work
Sonnet 5claude-sonnet-51M / 128k$2 / $10Best speed/intelligence balance; general default
Haiku 4.5claude-haiku-4-5200k / 64k$1 / $5Fastest/cheapest; high-volume, latency-sensitive, narrow tasks

Routing and cascades

text
┌──────────────┐ classify task difficulty / stakes
request ────▶ │ router │ (rules, or a Haiku classifier)
└──────┬───────┘
easy / high-vol │ hard / high-stakes
▼ ▼
┌────────────┐ ┌────────────┐
│ Haiku 4.5 │ │ Opus 5 / │
│ Sonnet 5 │──fail─▶│ Fable 5.1 │ (escalate on
└────────────┘ check └────────────┘ validation failure)
  • Static routing: rules by task type/segment (extraction → Haiku, synthesis → Sonnet, hardest coding → Opus/Fable).
  • Cascade: try cheap first; escalate on a validation/quality check failure, not on self-reported confidence.

Exam signal

“Cost is 4× budget but quality is fine” → introduce a cascade (cheap-first, escalate on failed checks). “We need the newest capability and cost is secondary” → Opus 5 / Fable 5.1. Distractors that escalate on self-reported confidence are wrong (anti-pattern #4).

Cost math worked example

100k requests/day, ~3k input + 1k output tokens each.

StrategyDaily cost (approx.)Note
All Opus 5100k × (3k×$5 + 1k×$25)/1e6 = 100k × $0.040 = $4,000baseline
All Sonnet 5100k × (3k×$2 + 1k×$10)/1e6 = 100k × $0.016 = $1,60060% cheaper
Cascade: 80% Haiku, 20% Opus80k×$0.008 + 20k×$0.040 = $1,440quality preserved on hard 20%
Sonnet + cache (80% cache hit on 2k prefix)~$960cache read ≈ 0.1× input

2.2 Managing breaking changes across versions

Model IDs from 4.6 onward are dateless but still pinned snapshots. Upgrading a model can change behaviour and can break code that relied on removed parameters.

ChangeAffectedMigration action
budget_tokens removed (returns 400)Fable 5.x, Opus 5, Sonnet 5, Opus 4.7–4.8Use thinking: {"type": "adaptive"} and effort; only Haiku 4.5 still uses budget_tokens
Forced tool use returns 400 on Fable 5.1Fable 5.1tool_choice:"any" / forced {"type":"tool"} → use auto + instruction, strict:true, or structured outputs
Thinking blocks readable only by producing model or newerall thinking modelsNever silently fall back to an older model mid-session
Opus 4.1 retired 2026-08-05pinned to 4.1Repin and re-run the eval/regression suite

Always re-run the regression suite (D4) before promoting a model change. Use client.models.list() / .retrieve(id) to confirm live limits.


2.3 Prompts and guardrails as governed assets

At Professional level a system prompt is a shared organisational asset, not a string in someone’s notebook. Govern it like code.

PropertyPractice
VersionedStored in VCS with semantic version; changes reviewed
TestableEach version runs against the golden eval set before release
ModularComposed of stable blocks (role, policy, format) + volatile blocks (task)
GuardrailedSafety and business rules layered, not buried in prose
Rolled outCanary → percentage → full, with rollback

Prompt-as-enforcement anti-pattern

Critical business rules (“never issue a refund over $500”) must be enforced programmatically (tool permission hooks, validation), not by a sentence in the system prompt. A model can be talked out of prose rules; a hook cannot. Options that rely on prompt wording to enforce a hard rule are wrong (anti-pattern #3).


2.4 Prompting techniques at the right altitude

TechniqueUse whenCost/latency note
Zero-shotTask is common and well-specifiedCheapest
Few-shotOutput shape/format must be pinned; edge cases shownAdds input tokens (cache the examples)
Chain-of-thoughtReasoning must be explicit (older/non-thinking models)More output tokens
Extended / adaptive thinkingHard reasoning; thinking:{"type":"adaptive"} on current modelsThinking tokens billed as output
Effort (low|medium|high|xhigh)Tune reasoning depth vs cost; xhigh for hardest coding/agentic on Opus 5 / Fable 5.1Higher effort = more tokens/latency
Fast modeLatency-sensitive useTrades some depth for speed

On current models, prefer adaptive thinking + effort over manual budgets. Only Haiku 4.5 still accepts budget_tokens and has **no effort parameter`.

json
{
"model": "claude-opus-5",
"thinking": { "type": "adaptive" },
"effort": "xhigh",
"messages": [{ "role": "user", "content": "Refactor this service for idempotency." }]
}

2.5 Context-window and token optimisation

A 1M-token window is a budget, not a target. Filling it raises cost and latency and can reduce accuracy (needle-in-haystack degradation).

LeverEffect
Retrieve, don’t stuffSend only relevant chunks (RAG, D3) instead of whole corpora
Prompt cachingAmortise stable prefix; cache read ≈ 0.1× input
Context editingClear stale tool results from the window
CompactionServer-side summarisation preserving narrative for long sessions
Output trimmingAsk for the schema you need; avoid verbose prose
Structured outputsFewer wasted tokens than free-form + reparse

2.6 Caching architecture: stable-prefix-first

Prompt caching only helps if the cache prefix is stable. Order content most-stable first: system prompt → tools → long documents → few-shot examples → volatile user turn. Mark the stable boundary with cache_control.

text
[ system prompt ] ← stable ┐
[ tool definitions ] ← stable │ cache_control: {"type":"ephemeral"} (this prefix is cached)
[ reference documents]← stable │
[ few-shot examples ] ← stable ┘
------------------------------------- cache boundary
[ user's actual question ] ← volatile (changes every request → never cache here)
  • Cache write ≈ 1.25× (5-min TTL) or 2× (1-hour TTL); cache read ≈ 0.1× base input.
  • Minimum cacheable prefix ~1024 tokens (2048 on Haiku).
  • Putting anything volatile before the stable content invalidates the cache on every call — a classic mistake.

Modular prompts + Skills: keep reusable capability blocks as Skills (SKILL.md, loaded progressively on demand) rather than pasting everything into every prompt. This keeps the cacheable prefix stable and the context lean.


2.7 Prompt versioning and rollout

Treat each prompt as name@semver. Store in VCS. Record which model version it was validated against — a prompt tuned for Opus 5 is not guaranteed to behave on Haiku 4.5. Tag the eval scores achieved.


2.8 Fable 5.1 append-only harness constraints

Fable 5.1 has thinking always on, and its thinking blocks are readable only by the producing model (or a newer one). Editing, reordering or removing earlier turns invalidates later thinking blocks. Therefore harnesses must be append-only.

RuleConsequence if violated
Freeze system and tools after the session startsEditing them invalidates downstream thinking
Put mid-session changes in a role: "system" message (append, don’t edit)Rewriting history breaks the loop
Trim server-side via context editing / compaction, not by deleting turns client-sideClient-side deletion invalidates thinking blocks
Never force tool use (tool_choice:"any" / forced tool) — returns 400Request fails; use auto + instruction / strict / structured outputs
Never silently fall back to an older modelOlder model drops the thinking blocks

Sonnet 5 differs

Sonnet 5 does not support mid-conversation system messages and has no task budgets. Do not assume Fable’s append-a-system-message trick works identically on Sonnet 5 — the exam tests these per-model differences.


2.9 Prompt-caching cost arithmetic

Caching only pays when you can quantify it. Read ≈ 0.1× base input; write ≈ 1.25× (5-min TTL) or 2× (1-hour TTL); minimum cacheable prefix ~1024 tokens (2048 on Haiku).

Worked example. Sonnet 5, 4,000-token stable prefix ($2/MTok input), 500-token volatile turn, 10,000 requests/day, 90% cache-hit rate after warm-up.

text
Uncached input cost/req = 4,500 × $2 / 1e6 = $0.0090
Cached (hit) input cost = (4,000 × 0.1 + 500) × $2/1e6 = (400 + 500)×$2/1e6 = $0.0018
Cache write (miss, 1.25×) = (4,000 × 1.25 + 500) × $2/1e6 ≈ $0.0110 (paid on ~10% of calls)
Daily uncached = 10,000 × $0.0090 = $90.00
Daily cached = 0.9×10,000×$0.0018 + 0.1×10,000×$0.0110 = $16.20 + $11.00 = $27.20
Saving ≈ 70% of input cost.

Exam signal

Caching helps in proportion to prefix size × hit rate. A tiny prefix or a low hit rate (because a volatile token sits in the prefix) makes caching worthless. If a stem says “hit rate is near zero”, look for a volatile element before the cache boundary — not “disable caching”.

SymptomCauseFix
Hit rate near zeroVolatile content before the boundary; per-request timestamp/user-id in prefixMove volatile content after the cache_control boundary
Write cost dominatesPrefix rarely reused within the TTLUse the 1-hour TTL, or don’t cache low-reuse prefixes
No effect on HaikuPrefix under the 2048-token minimumConsolidate stable context or accept no caching

2.10 Structured outputs and schema enforcement

At Professional level, “parse the prose” is never the answer. Use output_config.format with a JSON schema and strict: true so the model’s output conforms by construction, and reserve validation-retry for the rare miss.

json
{
"model": "claude-sonnet-5",
"messages": [{ "role": "user", "content": "Extract the invoice fields." }],
"output_config": {
"format": {
"type": "json_schema",
"schema": {
"type": "object",
"properties": {
"invoice_id": { "type": "string" },
"total": { "type": "number" },
"currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] }
},
"required": ["invoice_id", "total", "currency"],
"additionalProperties": false
},
"strict": true
}
}
}
ApproachReliabilityWhen
Free-text + regex/parseBrittleNever for structured data
Prompt “return JSON” onlyBetter, still fallibleLegacy/unsupported paths
strict JSON schema (structured outputs)Conforms by constructionDefault for machine-consumed output
Tool schema with strict: trueEnforced tool argumentsWhen a tool needs typed args

Exam signal

On Fable 5.1 you cannot force tool use (tool_choice:"any" → 400). To guarantee a shape, use structured outputs / strict schema or auto + instruction — the exam pairs the “guaranteed JSON” need with the “no forced tools on Fable” constraint.


2.11 Context editing vs compaction

Long-running sessions overflow the window. Two server-side tools manage it, and they are not interchangeable.

TechniqueWhat it doesUse whenRisk if misused
Context editingRemoves/clears stale tool results and blocks from the windowTool outputs are large and no longer neededEditing earlier turns invalidates Fable 5.1 thinking blocks — edit tool results, not reasoning
CompactionServer-side summarisation preserving the narrativeVery long sessions where history must be retained in gistOver-compaction loses detail needed later
Memory toolPersist durable facts outside the windowFacts must survive across sessionsStoring secrets/PII inappropriately

Prefer trimming server-side (context editing / compaction) over client-side deletion of turns, which breaks append-only harness invariants (2.8). The PreCompact hook lets you snapshot state before compaction runs.


2.12 Scenario walkthrough: taming a runaway prompt-and-model bill

Scenario. A document-analysis product runs every request on Opus 5 with effort: xhigh, a 9k-token system prompt duplicated per call, and forced tool use. Monthly spend is 5× budget; the team also just failed a Fable 5.1 pilot with 400 errors. Quality is acceptable; latency is not the complaint — cost is.

Expert reasoning trace.

  1. Right-size the model. Quality is already acceptable on Opus 5, so most traffic can run on Sonnet 5 with a cascade escalating to Opus 5 only on a validation-check failure. That alone cuts the per-request rate ~60%.

  2. Fix the effort. xhigh everywhere is wasteful; drop to high/medium and re-run the regression suite per segment to confirm no drop.

  3. Cache the prefix. The 9k-token system prompt is stable → mark a cache_control boundary; move the per-request document after it. Reads at 0.1× turn the prefix nearly free at a high hit rate.

  4. Modularise with Skills. The duplicated capability text belongs in Skills loaded on demand, keeping the cached prefix lean and stable.

  5. Explain the Fable 400s. Forced tool use is unsupported on Fable 5.1; switch to auto + instruction or structured outputs. But note Fable’s $10/$50 pricing makes it the wrong cost choice here anyway.

  6. Re-validate and roll out. Offline regression per segment → canary → ramp, prior version hot for rollback.

Why the tempting alternatives are wrong: “move everything to Haiku” risks the quality bar; “escalate on self-reported confidence” is anti-pattern #4; “just buy a bigger budget” ignores the arithmetic; “keep forcing tools and retry” cannot fix a 400.


2.13 Common misconceptions

MisconceptionRealityWhy it matters on the exam
“budget_tokens is the standard way to control thinking.”Removed on current models (400); only Haiku 4.5 still uses it. Use adaptive thinking + effort.A 400-after-upgrade stem tests exactly this.
“A strong system-prompt sentence enforces a business rule.”Prompts are guidance; hard rules need hooks/validation.Prompt-as-enforcement is a recurring wrong answer.
“Higher effort always means better answers.”Beyond the task’s need it just adds tokens/latency/cost.xhigh-everywhere distractors overspend.
“Caching automatically saves money once enabled.”Only if the prefix is stable and reused above the minimum size.Volatile-prefix stems make caching worthless.
“Forcing tool use works on every model.”Fable 5.1 returns 400 on forced tool use.Guaranteed-shape stems pair with structured outputs.
“A prompt tuned on Opus behaves the same on Haiku.”Behaviour differs per model; re-validate per model.Cross-model reuse without re-eval is the trap.
“You can trim a long session by deleting old turns client-side.”On thinking models that invalidates later thinking blocks; trim server-side.Append-only harness rules are tested per model.

Exam traps in this domain

TrapWhy it is wrong
Using the most expensive model for every requestIgnores routing/cascades; blows the cost budget
Escalating in a cascade on self-reported confidenceSelf-report is unreliable (anti-pattern #4); escalate on validation failure
Enforcing a hard business rule via the system promptPrompt-as-enforcement (anti-pattern #3); use hooks/validation
Setting budget_tokens on Opus 5 / Sonnet 5 / Fable 5.1Removed; returns 400 — use adaptive thinking + effort
Forcing tool use on Fable 5.1Returns 400; use auto+instruction, strict, or structured outputs
Editing earlier turns in a Fable 5.1 sessionInvalidates later thinking blocks; harness must be append-only
Silently falling back to an older model mid-sessionDrops thinking blocks; corrupts the session
Putting the volatile user turn before the cached prefixInvalidates the cache every call
Filling the 1M window “because it’s available”Raises cost/latency; can reduce accuracy; retrieve instead
Assuming a prompt tuned on one model behaves identically on anotherMust re-validate per model version
Parsing free-text prose for structured data instead of using a strict schemaBrittle; structured outputs conform by construction
Deleting old turns client-side to trim a thinking-model sessionInvalidates later thinking blocks; use context editing/compaction
Assuming caching saves money regardless of prefix size or hit rateSaving ∝ prefix size × hit rate; a volatile prefix yields ~0
Storing secrets/PII in the memory tool or persisted contextExfiltration/compliance risk; keep secrets in secret managers
Raising effort to xhigh to “improve quality” without evidenceAdds tokens/latency/cost past the task’s need

Practice questions

Q1 · A pipeline routes everything to Opus 5. Cost is 4× budget; quality is acceptable. Which change best cuts cost while preserving quality on hard cases? (Select one)

A. Move all traffic to Haiku 4.5. B. Build a cascade: Haiku 4.5 / Sonnet 5 first, escalate to Opus 5 only when an output validation check fails, and cache the stable prefix. C. Escalate to Opus 5 whenever the model reports low confidence in its own answer. D. Increase effort to xhigh everywhere.

Answer: B. Cheap-first with escalation on validation failure preserves quality on hard cases while most traffic runs cheaply; caching amortises the stable prefix. Blanket Haiku (A) sacrifices quality. Self-reported confidence (C) is an anti-pattern. Raising effort everywhere (D) increases cost.

Q2 · A team upgrades from Opus 4.6 to Opus 5 and their requests now return 400 errors. The requests set `budget_tokens` for thinking. What is the fix? (Select one)

A. Add more retries. B. Replace budget_tokens with thinking: {\"type\": \"adaptive\"} and control depth via effort, since budget_tokens is removed on Opus 5. C. Downgrade permanently to Haiku 4.5. D. Remove thinking entirely.

Answer: B. budget_tokens is removed on Opus 5 (and Sonnet 5 / Fable 5.x) and returns 400; adaptive thinking plus effort is the supported replacement. Retries (A) won’t fix a 400. Haiku (C) is the only model still using budget_tokens but is not an equivalent for Opus workloads. Removing thinking (D) discards needed reasoning.

Q3 · A refund agent must never issue refunds above $500. Where should this rule live? (Select one)

A. As a firmly worded sentence in the system prompt. B. As a programmatic tool-permission hook / validation that rejects any refund over $500 before execution. C. As a few-shot example showing a refused large refund. D. In the model’s thinking budget.

Answer: B. Hard business rules require programmatic enforcement — a hook or validation the model cannot talk its way past. Prompt wording (A) and few-shot examples (C) are prompt-as-enforcement anti-patterns. Thinking budget (D) is unrelated.

Q4 · Prompt caching is enabled but hit rate is near zero. The prompt places the user's question first, then the system prompt and reference documents. Why, and what fixes it? (Select one)

A. Caching is broken; disable it. B. The volatile user turn sits before the stable content, so the cached prefix changes every call — reorder to stable-first (system → tools → docs) then the user turn, and mark the stable boundary with cache_control. C. The documents are too short. D. Haiku doesn’t support caching.

Answer: B. Caching keys on a stable prefix; putting the changing user turn first invalidates it every call. Stable-prefix-first with a cache_control boundary fixes it. Caching is not broken (A); document length (C) matters only for the ~1024-token minimum; Haiku does support caching (D) with a 2048-token minimum.

Q5 · Which are valid reasons NOT to stuff a whole 500k-token corpus into the 1M window every request? (Select two)

A. Higher token cost and latency per call. B. Possible accuracy degradation locating the relevant needle. C. The window physically cannot hold it. D. Structured outputs are disabled above 200k tokens. E. Caching is prohibited on large inputs.

Answer: A and B. Stuffing raises cost and latency and can hurt retrieval accuracy within a huge context; retrieval (RAG) sends only relevant chunks. It fits the window (C is false at 500k of 1M), structured outputs are not size-gated that way (D), and caching is allowed on large inputs (E).

Q6 · In an active Fable 5.1 agentic session, the team wants to change the system instructions mid-run. What is the correct approach? (Select one)

A. Edit the original system field in place. B. Append the change as a new role: \"system\" message, leaving earlier turns untouched, because the harness must be append-only. C. Delete the earliest turns to make room. D. Reorder messages to put the new instruction first.

Answer: B. Fable 5.1 thinking blocks are invalidated by editing/reordering/removing earlier turns, so mid-session changes are appended as a new system message. Editing in place (A), deleting turns (C) and reordering (D) all invalidate downstream thinking blocks.

Q7 · A team wants to force Fable 5.1 to always return a tool call using tool_choice set to any. It returns 400. What should they do? (Select one)

A. Retry until it works. B. Use tool_choice: \"auto\" with an instruction to use the tool, set strict: true on the tool schema, or use structured outputs — because forced tool use is unsupported on Fable 5.1. C. Switch to Haiku 4.5 permanently. D. Remove all tools.

Answer: B. Forced tool use (any / forced tool) returns 400 on Fable 5.1; the supported paths are auto + instruction, strict schemas, or structured outputs. Retrying (A) won’t fix a 400. Switching model (C) or removing tools (D) abandons the requirement.

Q8 · How should a new system-prompt version be rolled out to production? (Select one)

A. Swap it org-wide immediately to move fast. B. Validate offline against the golden set, canary on a small traffic slice with regression guardrails, ramp by percentage, and keep the prior version hot for instant rollback. C. Let each engineer edit the inline prompt in their own service. D. Ship it and monitor customer complaints.

Answer: B. Prompts are governed, versioned assets: offline eval → canary → ramp → keep prior version for rollback. Org-wide swaps (A) and complaint-driven monitoring (D) skip validation. Per-engineer inline edits (C) destroy governance and reproducibility.

Q9 · A high-volume extraction subtask feeds a slower synthesis step. Which portfolio assignment is BEST? (Select one)

A. Opus 5 for both steps. B. Haiku 4.5 for the high-volume extraction; Sonnet 5 or Opus 5 for the synthesis — matching model cost to each step’s difficulty. C. Fable 5.1 for both, for maximum quality. D. Haiku 4.5 for both, for maximum savings.

Answer: B. Portfolio routing assigns the cheapest adequate model per step: Haiku for narrow high-volume extraction, a stronger model for harder synthesis. Opus/Fable for both (A, C) overpays; Haiku for both (D) risks the synthesis quality.

Q10 · Reusable capability blocks are being pasted into every prompt, bloating context and breaking the cache prefix. What is the better pattern? (Select one)

A. Package them as Skills (SKILL.md) loaded progressively on demand, keeping the stable cacheable prefix lean. B. Duplicate them into each service’s prompt. C. Move them into the volatile user turn. D. Increase the context window.

Answer: A. Skills load capability progressively on demand, keeping context lean and the cache prefix stable. Duplication (B) is what caused the bloat; moving them to the volatile turn (C) worsens caching; a bigger window (D) doesn’t address cost or cache stability.

Q11 · Which statement about current-model thinking configuration is correct? (Select one)

A. All current models require budget_tokens. B. Current models use thinking: {\"type\":\"adaptive\"} with effort levels low/medium/high/xhigh; only Haiku 4.5 still uses budget_tokens and has no effort parameter. C. Effort only exists on Haiku 4.5. D. xhigh effort is the default on all models.

Answer: B. Adaptive thinking plus effort is standard on current models; Haiku 4.5 is the exception still using budget_tokens and lacking effort. budget_tokens is not universal (A); effort is not Haiku-only (C); high (not xhigh) is the default (D).

Q12 · A prompt validated on Opus 5 is reused verbatim on Haiku 4.5 and quality drops. What is the correct lesson? (Select one)

A. Haiku 4.5 is defective. B. Prompts are validated per model version; a prompt tuned for one model must be re-tested (and often adjusted) against the golden set on any other model before use. C. Always use Opus 5. D. Quality drops are unavoidable and should be ignored.

Answer: B. Model behaviour differs across the portfolio, so prompt versions carry the model they were validated against and must be re-evaluated when reused elsewhere. Haiku is not defective (A); mandating Opus (C) ignores cost; ignoring regressions (D) is negligent.

Q13 · A 4,000-token stable prefix on Sonnet 5 is reused with a 90% cache-hit rate at 10,000 req/day. Roughly what does caching save on input cost? (Select one)

A. Nothing; caching never helps large prefixes. B. Around 70%, because a cache read is ~0.1× input so the 4k prefix becomes ~400 effective tokens on hits. C. Exactly 50%, the Batch discount. D. 100%; cached requests are free.

Answer: B. Read ≈ 0.1× input turns 4,000 prefix tokens into ~400 on the 90% of hits; the daily input cost drops from ~$90 to ~$27, about 70%. Caching does help large stable prefixes (A); 50% is the Batch discount, not caching (C); cached reads are cheap, not free (D).

Q14 · A pipeline must return a strictly typed JSON object for downstream systems, and it runs on Fable 5.1 where forcing tool use returns 400. What is the BEST approach? (Select one)

A. Force a tool call with tool_choice: 'any' and retry on 400. B. Use structured outputs with a strict JSON schema (or auto + instruction), which guarantees the shape without forcing tool use. C. Ask for JSON in the prompt and regex-parse the prose. D. Switch to free-text and reparse.

Answer: B. Structured outputs with a strict schema conform by construction and don’t require forced tool use, which Fable 5.1 rejects with 400. Forcing tools (A) fails; prompt-only JSON with regex (C) and free-text reparse (D) are brittle prose-parsing.

Q15 · A long agentic session on a thinking model overflows the window because tool results are huge. Which technique is correct, and what must be avoided? (Select one)

A. Delete the earliest user/assistant turns client-side. B. Use context editing to clear stale tool results server-side (and compaction for narrative), avoiding client-side edits to earlier reasoning that invalidate thinking blocks. C. Lower the temperature to shrink the context. D. Force the model to summarise itself in the same turn.

Answer: B. Context editing removes stale tool results server-side; compaction summarises narrative — both avoid invalidating thinking blocks. Deleting turns client-side (A) breaks the append-only invariant; temperature (C) doesn’t affect context size; in-turn self-summary (D) doesn’t reclaim the window.

Q16 · A team runs every request on Opus 5 at `effort: xhigh` with acceptable quality and 5× budget; cost, not latency, is the complaint. Which TWO changes best cut cost while preserving quality? (Select two)

A. Cascade most traffic to Sonnet 5, escalating to Opus 5 only on a validation-check failure. B. Lower effort to an adequate level and re-run the per-segment regression suite. C. Escalate on the model’s self-reported confidence. D. Move all traffic to Fable 5.1 for quality. E. Remove the eval suite to cut compute.

Answer: A and B. A cheap-first cascade and right-sized effort (re-validated per segment) cut cost while preserving quality on hard cases. Self-report (C) is anti-pattern #4; Fable 5.1 (D) is the most expensive model; removing evals (E) removes the quality guard.

Q17 · Caching is enabled but the hit rate is ~3%. Investigation shows a per-request `request_id` string is prepended to the system prompt. What is the fix? (Select one)

A. Disable caching; it doesn’t work here. B. Remove the volatile request_id from the prefix (log it separately) so the prefix is byte-stable, restoring cache hits. C. Shorten the documents. D. Increase the context window.

Answer: B. A per-request token in the prefix changes it every call, so nothing caches; moving it out restores a stable prefix. Caching isn’t broken (A); document length (C) only affects the minimum; a bigger window (D) is unrelated.

Q18 · A durable fact (a customer's contract tier) must persist across separate sessions without re-sending it in every prompt. Which mechanism fits, and what constraint applies? (Select one)

A. Paste the fact into every system prompt. B. Use the memory tool to persist the fact across sessions, but never store secrets/PII there inappropriately and keep it out of model-visible logs. C. Store it in CLAUDE.local.md. D. Increase retention to keep it in provider logs.

Answer: B. The memory tool persists durable facts across sessions; the constraint is not to store secrets/PII inappropriately. Pasting per prompt (A) bloats context; CLAUDE.local.md (C) is a Claude Code dev file, not a runtime store; relying on provider retention (D) is not a memory mechanism and raises compliance risk.

Key takeaways

  • Manage a portfolio: route/cascade to the cheapest model that clears the quality bar; escalate on validation failure, not self-reported confidence.
  • Do the cost math — cascades and caching routinely cut spend 60–75% with quality preserved on hard cases.
  • Manage breaking changes: budget_tokens removed (400) except on Haiku 4.5; forced tool use fails on Fable 5.1; re-run regressions before promoting a model.
  • Treat prompts and guardrails as versioned, governed assets; enforce hard rules programmatically, never via prompt prose.
  • Use adaptive thinking + effort; reserve xhigh for the hardest Opus/Fable work.
  • Cache stable-prefix-first; keep reusable blocks as Skills; never place volatile content before the cached prefix.
  • Fable 5.1 harnesses are append-only: freeze system/tools, append system messages, trim server-side, never force tools or silently downgrade.
  • Caching saving ∝ prefix size × hit rate; a volatile token in the prefix (timestamps, request IDs) drops the hit rate to ~0 — remove it, don’t disable caching.
  • Guarantee output shape with structured outputs / strict schemas, especially on Fable 5.1 where forced tool use returns 400 — never parse prose.
  • Manage long sessions with context editing (clear stale tool results) and compaction (summarise narrative) server-side; the memory tool persists durable facts (no secrets/PII).

Last updated Sep 18, 2026