# D2 · Claude Models, Prompting & Context Engineering

Portfolio-level model selection and routing, prompts and guardrails as governed assets, thinking/effort control, context and token optimisation, caching architecture, prompt versioning, and Fable 5.1 harness constraints.

import { Accordions, AccordionItem, Tabs, TabItem, Steps } from '@prosefly/astro-components';

This domain is worth **13% – roughly 8 of 63 items**. At Professional level you are not writing a single prompt; you are managing a **portfolio of models and a library of governed prompts** across an organisation, with cost math, versioning, rollout and breaking-change management. Items reward the option that treats prompts, models and thinking configuration as **engineered, versioned assets** with measurable trade-offs.

## Learning objectives

By the end of this page you should be able to:

1. Select and **route/cascade** across the model portfolio (Fable 5.1, Opus 5, Sonnet 5, Haiku 4.5) with real cost math.
2. Manage **breaking changes** across model versions (thinking-block rules, removed parameters).
3. Treat **system prompts, templates and guardrails** as governed, versioned assets.
4. Apply prompting techniques (**zero-shot, few-shot, CoT, extended/adaptive thinking, effort**) at the right altitude.
5. Optimise **context window and tokens**.
6. Design a **caching architecture** (stable-prefix-first, modular prompts, Skills) and a **versioning/rollout** process.
7. Respect **Fable 5.1's append-only harness constraints**.

---

## 2.1 The model portfolio and routing

There is no single "best" model — there is a portfolio, and the architect's job is to route each request to the cheapest model that meets its quality bar.

| Model | ID | Context / max out | In / out per MTok | When it is the right default |
| --- | --- | --- | --- | --- |
| Fable 5.1 | `claude-fable-5-1` | 1M / 128k | \$10 / \$50 | Frontier reasoning; thinking always on; append-only harness; 30-day retention |
| Opus 5 | `claude-opus-5` | 1M / 128k | \$5 / \$25 | Complex agentic coding and enterprise work |
| Sonnet 5 | `claude-sonnet-5` | 1M / 128k | \$2 / \$10 | Best speed/intelligence balance; general default |
| Haiku 4.5 | `claude-haiku-4-5` | 200k / 64k | \$1 / \$5 | Fastest/cheapest; high-volume, latency-sensitive, narrow tasks |

### Routing and cascades

```text
                 ┌──────────────┐  classify task difficulty / stakes
   request ────▶ │   router     │  (rules, or a Haiku classifier)
                 └──────┬───────┘
        easy / high-vol │        hard / high-stakes
                        ▼                    ▼
                 ┌────────────┐       ┌────────────┐
                 │ Haiku 4.5  │       │  Opus 5 /   │
                 │ Sonnet 5   │──fail─▶│  Fable 5.1 │  (escalate on
                 └────────────┘  check └────────────┘   validation failure)
```

- **Static routing**: rules by task type/segment (extraction → Haiku, synthesis → Sonnet, hardest coding → Opus/Fable).
- **Cascade**: try cheap first; escalate on a **validation/quality check** failure, not on self-reported confidence.

:::tip[Exam signal]
"Cost is 4× budget but quality is fine" → introduce a **cascade** (cheap-first, escalate on failed checks). "We need the newest capability and cost is secondary" → Opus 5 / Fable 5.1. Distractors that escalate on **self-reported confidence** are wrong (anti-pattern #4).
:::

### Cost math worked example

100k requests/day, ~3k input + 1k output tokens each.

| Strategy | Daily cost (approx.) | Note |
| --- | --- | --- |
| All Opus 5 | 100k × (3k×\$5 + 1k×\$25)/1e6 = 100k × \$0.040 = **\$4,000** | baseline |
| All Sonnet 5 | 100k × (3k×\$2 + 1k×\$10)/1e6 = 100k × \$0.016 = **\$1,600** | 60% cheaper |
| Cascade: 80% Haiku, 20% Opus | 80k×\$0.008 + 20k×\$0.040 = **\$1,440** | quality preserved on hard 20% |
| Sonnet + cache (80% cache hit on 2k prefix) | ~**\$960** | cache read ≈ 0.1× input |

---

## 2.2 Managing breaking changes across versions

Model IDs from 4.6 onward are **dateless but still pinned snapshots**. Upgrading a model can change behaviour and can break code that relied on removed parameters.

| Change | Affected | Migration action |
| --- | --- | --- |
| `budget_tokens` removed (returns 400) | Fable 5.x, Opus 5, Sonnet 5, Opus 4.7–4.8 | Use `thinking: {"type": "adaptive"}` and `effort`; only Haiku 4.5 still uses `budget_tokens` |
| Forced tool use returns 400 on Fable 5.1 | Fable 5.1 | `tool_choice:"any"` / forced `{"type":"tool"}` → use `auto` + instruction, `strict:true`, or structured outputs |
| Thinking blocks readable only by producing model or newer | all thinking models | Never silently fall back to an older model mid-session |
| Opus 4.1 retired 2026-08-05 | pinned to 4.1 | Repin and re-run the eval/regression suite |

Always re-run the **regression suite** (D4) before promoting a model change. Use `client.models.list()` / `.retrieve(id)` to confirm live limits.

---

## 2.3 Prompts and guardrails as governed assets

At Professional level a system prompt is a **shared organisational asset**, not a string in someone's notebook. Govern it like code.

| Property | Practice |
| --- | --- |
| Versioned | Stored in VCS with semantic version; changes reviewed |
| Testable | Each version runs against the golden eval set before release |
| Modular | Composed of stable blocks (role, policy, format) + volatile blocks (task) |
| Guardrailed | Safety and business rules layered, not buried in prose |
| Rolled out | Canary → percentage → full, with rollback |

:::caution[Prompt-as-enforcement anti-pattern]
Critical business rules ("never issue a refund over \$500") must be enforced **programmatically** (tool permission hooks, validation), not by a sentence in the system prompt. A model can be talked out of prose rules; a hook cannot. Options that rely on prompt wording to enforce a hard rule are wrong (anti-pattern #3).
:::

---

## 2.4 Prompting techniques at the right altitude

| Technique | Use when | Cost/latency note |
| --- | --- | --- |
| Zero-shot | Task is common and well-specified | Cheapest |
| Few-shot | Output shape/format must be pinned; edge cases shown | Adds input tokens (cache the examples) |
| Chain-of-thought | Reasoning must be explicit (older/non-thinking models) | More output tokens |
| Extended / adaptive thinking | Hard reasoning; `thinking:{"type":"adaptive"}` on current models | Thinking tokens billed as output |
| Effort (`low\|medium\|high\|xhigh`) | Tune reasoning depth vs cost; `xhigh` for hardest coding/agentic on Opus 5 / Fable 5.1 | Higher effort = more tokens/latency |
| Fast mode | Latency-sensitive use | Trades some depth for speed |

On current models, prefer **adaptive thinking + effort** over manual budgets. Only **Haiku 4.5** still accepts `budget_tokens` and has **no `effort` parameter`.

```json
{
  "model": "claude-opus-5",
  "thinking": { "type": "adaptive" },
  "effort": "xhigh",
  "messages": [{ "role": "user", "content": "Refactor this service for idempotency." }]
}
```

---

## 2.5 Context-window and token optimisation

A 1M-token window is a budget, not a target. Filling it raises cost and latency and can *reduce* accuracy (needle-in-haystack degradation).

| Lever | Effect |
| --- | --- |
| Retrieve, don't stuff | Send only relevant chunks (RAG, D3) instead of whole corpora |
| Prompt caching | Amortise stable prefix; cache read ≈ 0.1× input |
| Context editing | Clear stale tool results from the window |
| Compaction | Server-side summarisation preserving narrative for long sessions |
| Output trimming | Ask for the schema you need; avoid verbose prose |
| Structured outputs | Fewer wasted tokens than free-form + reparse |

---

## 2.6 Caching architecture: stable-prefix-first

Prompt caching only helps if the **cache prefix is stable**. Order content **most-stable first**: system prompt → tools → long documents → few-shot examples → volatile user turn. Mark the stable boundary with `cache_control`.

```text
[ system prompt ]      ← stable  ┐
[ tool definitions ]   ← stable  │  cache_control: {"type":"ephemeral"}  (this prefix is cached)
[ reference documents]← stable  │
[ few-shot examples ]  ← stable  ┘
------------------------------------- cache boundary
[ user's actual question ]  ← volatile (changes every request → never cache here)
```

- Cache **write** ≈ 1.25× (5-min TTL) or 2× (1-hour TTL); cache **read** ≈ 0.1× base input.
- Minimum cacheable prefix ~**1024 tokens** (2048 on Haiku).
- Putting anything volatile before the stable content **invalidates the cache** on every call — a classic mistake.

**Modular prompts + Skills**: keep reusable capability blocks as **Skills** (`SKILL.md`, loaded progressively on demand) rather than pasting everything into every prompt. This keeps the cacheable prefix stable and the context lean.

---

## 2.7 Prompt versioning and rollout

<Tabs>
  <TabItem label="Versioning">
    Treat each prompt as `name@semver`. Store in VCS. Record which model version it was validated against — a prompt tuned for Opus 5 is not guaranteed to behave on Haiku 4.5. Tag the eval scores achieved.
  </TabItem>
  <TabItem label="Rollout">
    Canary the new version on a small traffic slice with online metrics and a guardrail on regression. Ramp by percentage. Keep the previous version hot for instant rollback. Never swap a prompt org-wide without an offline eval pass first.
  </TabItem>
  <TabItem label="Rollback">
    Because prompts are versioned and the prior version stays deployable, rollback is a config flip, not a redeploy. This is why prompts belong in a governed registry, not inline in application code.
  </TabItem>
</Tabs>

---

## 2.8 Fable 5.1 append-only harness constraints

Fable 5.1 has **thinking always on**, and its **thinking blocks are readable only by the producing model (or a newer one)**. Editing, reordering or removing earlier turns **invalidates later thinking blocks**. Therefore harnesses must be **append-only**.

| Rule | Consequence if violated |
| --- | --- |
| Freeze `system` and `tools` after the session starts | Editing them invalidates downstream thinking |
| Put mid-session changes in a `role: "system"` message (append, don't edit) | Rewriting history breaks the loop |
| Trim server-side via **context editing / compaction**, not by deleting turns client-side | Client-side deletion invalidates thinking blocks |
| Never force tool use (`tool_choice:"any"` / forced tool) — returns 400 | Request fails; use `auto` + instruction / `strict` / structured outputs |
| Never silently fall back to an older model | Older model drops the thinking blocks |

:::note[Sonnet 5 differs]
Sonnet 5 does **not** support mid-conversation system messages and has no task budgets. Do not assume Fable's append-a-system-message trick works identically on Sonnet 5 — the exam tests these per-model differences.
:::

---

## 2.9 Prompt-caching cost arithmetic

Caching only pays when you can quantify it. Read ≈ **0.1×** base input; write ≈ **1.25×** (5-min TTL) or **2×** (1-hour TTL); minimum cacheable prefix ~**1024** tokens (2048 on Haiku).

**Worked example.** Sonnet 5, 4,000-token stable prefix (\$2/MTok input), 500-token volatile turn, 10,000 requests/day, 90% cache-hit rate after warm-up.

```text
Uncached input cost/req  = 4,500 × $2 / 1e6            = $0.0090
Cached (hit) input cost  = (4,000 × 0.1 + 500) × $2/1e6 = (400 + 500)×$2/1e6 = $0.0018
Cache write (miss, 1.25×) = (4,000 × 1.25 + 500) × $2/1e6 ≈ $0.0110  (paid on ~10% of calls)

Daily uncached  = 10,000 × $0.0090                    = $90.00
Daily cached    = 0.9×10,000×$0.0018 + 0.1×10,000×$0.0110 = $16.20 + $11.00 = $27.20
Saving ≈ 70% of input cost.
```

:::tip[Exam signal]
Caching helps in proportion to **prefix size × hit rate**. A tiny prefix or a low hit rate (because a volatile token sits in the prefix) makes caching worthless. If a stem says "hit rate is near zero", look for a volatile element before the cache boundary — not "disable caching".
:::

| Symptom | Cause | Fix |
| --- | --- | --- |
| Hit rate near zero | Volatile content before the boundary; per-request timestamp/user-id in prefix | Move volatile content after the `cache_control` boundary |
| Write cost dominates | Prefix rarely reused within the TTL | Use the 1-hour TTL, or don't cache low-reuse prefixes |
| No effect on Haiku | Prefix under the 2048-token minimum | Consolidate stable context or accept no caching |

---

## 2.10 Structured outputs and schema enforcement

At Professional level, "parse the prose" is never the answer. Use `output_config.format` with a JSON schema and `strict: true` so the model's output conforms by construction, and reserve validation-retry for the rare miss.

```json
{
  "model": "claude-sonnet-5",
  "messages": [{ "role": "user", "content": "Extract the invoice fields." }],
  "output_config": {
    "format": {
      "type": "json_schema",
      "schema": {
        "type": "object",
        "properties": {
          "invoice_id": { "type": "string" },
          "total": { "type": "number" },
          "currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] }
        },
        "required": ["invoice_id", "total", "currency"],
        "additionalProperties": false
      },
      "strict": true
    }
  }
}
```

| Approach | Reliability | When |
| --- | --- | --- |
| Free-text + regex/parse | Brittle | Never for structured data |
| Prompt "return JSON" only | Better, still fallible | Legacy/unsupported paths |
| `strict` JSON schema (structured outputs) | Conforms by construction | Default for machine-consumed output |
| Tool schema with `strict: true` | Enforced tool arguments | When a tool needs typed args |

:::tip[Exam signal]
On **Fable 5.1** you cannot force tool use (`tool_choice:"any"` → 400). To guarantee a shape, use **structured outputs / `strict` schema** or `auto` + instruction — the exam pairs the "guaranteed JSON" need with the "no forced tools on Fable" constraint.
:::

---

## 2.11 Context editing vs compaction

Long-running sessions overflow the window. Two server-side tools manage it, and they are not interchangeable.

| Technique | What it does | Use when | Risk if misused |
| --- | --- | --- | --- |
| **Context editing** | Removes/clears stale tool results and blocks from the window | Tool outputs are large and no longer needed | Editing earlier turns invalidates Fable 5.1 thinking blocks — edit *tool results*, not reasoning |
| **Compaction** | Server-side summarisation preserving the narrative | Very long sessions where history must be retained in gist | Over-compaction loses detail needed later |
| **Memory tool** | Persist durable facts outside the window | Facts must survive across sessions | Storing secrets/PII inappropriately |

Prefer trimming **server-side** (context editing / compaction) over client-side deletion of turns, which breaks append-only harness invariants (2.8). The `PreCompact` hook lets you snapshot state before compaction runs.

---

## 2.12 Scenario walkthrough: taming a runaway prompt-and-model bill

**Scenario.** A document-analysis product runs every request on Opus 5 with `effort: xhigh`, a 9k-token system prompt duplicated per call, and forced tool use. Monthly spend is 5× budget; the team also just failed a Fable 5.1 pilot with 400 errors. Quality is acceptable; latency is not the complaint — cost is.

**Expert reasoning trace.**

<Steps>

1. **Right-size the model.** Quality is already acceptable on Opus 5, so most traffic can run on **Sonnet 5** with a cascade escalating to Opus 5 only on a validation-check failure. That alone cuts the per-request rate ~60%.

2. **Fix the effort.** `xhigh` everywhere is wasteful; drop to `high`/`medium` and re-run the regression suite per segment to confirm no drop.

3. **Cache the prefix.** The 9k-token system prompt is stable → mark a `cache_control` boundary; move the per-request document *after* it. Reads at 0.1× turn the prefix nearly free at a high hit rate.

4. **Modularise with Skills.** The duplicated capability text belongs in **Skills** loaded on demand, keeping the cached prefix lean and stable.

5. **Explain the Fable 400s.** Forced tool use is unsupported on Fable 5.1; switch to `auto` + instruction or **structured outputs**. But note Fable's \$10/\$50 pricing makes it the *wrong* cost choice here anyway.

6. **Re-validate and roll out.** Offline regression per segment → canary → ramp, prior version hot for rollback.

</Steps>

**Why the tempting alternatives are wrong:** "move everything to Haiku" risks the quality bar; "escalate on self-reported confidence" is anti-pattern #4; "just buy a bigger budget" ignores the arithmetic; "keep forcing tools and retry" cannot fix a 400.

---

## 2.13 Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "`budget_tokens` is the standard way to control thinking." | Removed on current models (400); only Haiku 4.5 still uses it. Use adaptive thinking + `effort`. | A 400-after-upgrade stem tests exactly this. |
| "A strong system-prompt sentence enforces a business rule." | Prompts are guidance; hard rules need hooks/validation. | Prompt-as-enforcement is a recurring wrong answer. |
| "Higher effort always means better answers." | Beyond the task's need it just adds tokens/latency/cost. | `xhigh`-everywhere distractors overspend. |
| "Caching automatically saves money once enabled." | Only if the prefix is stable and reused above the minimum size. | Volatile-prefix stems make caching worthless. |
| "Forcing tool use works on every model." | Fable 5.1 returns 400 on forced tool use. | Guaranteed-shape stems pair with structured outputs. |
| "A prompt tuned on Opus behaves the same on Haiku." | Behaviour differs per model; re-validate per model. | Cross-model reuse without re-eval is the trap. |
| "You can trim a long session by deleting old turns client-side." | On thinking models that invalidates later thinking blocks; trim server-side. | Append-only harness rules are tested per model. |

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| Using the most expensive model for every request | Ignores routing/cascades; blows the cost budget |
| Escalating in a cascade on self-reported confidence | Self-report is unreliable (anti-pattern #4); escalate on validation failure |
| Enforcing a hard business rule via the system prompt | Prompt-as-enforcement (anti-pattern #3); use hooks/validation |
| Setting `budget_tokens` on Opus 5 / Sonnet 5 / Fable 5.1 | Removed; returns 400 — use adaptive thinking + effort |
| Forcing tool use on Fable 5.1 | Returns 400; use `auto`+instruction, `strict`, or structured outputs |
| Editing earlier turns in a Fable 5.1 session | Invalidates later thinking blocks; harness must be append-only |
| Silently falling back to an older model mid-session | Drops thinking blocks; corrupts the session |
| Putting the volatile user turn before the cached prefix | Invalidates the cache every call |
| Filling the 1M window "because it's available" | Raises cost/latency; can reduce accuracy; retrieve instead |
| Assuming a prompt tuned on one model behaves identically on another | Must re-validate per model version |
| Parsing free-text prose for structured data instead of using a `strict` schema | Brittle; structured outputs conform by construction |
| Deleting old turns client-side to trim a thinking-model session | Invalidates later thinking blocks; use context editing/compaction |
| Assuming caching saves money regardless of prefix size or hit rate | Saving ∝ prefix size × hit rate; a volatile prefix yields ~0 |
| Storing secrets/PII in the memory tool or persisted context | Exfiltration/compliance risk; keep secrets in secret managers |
| Raising `effort` to `xhigh` to "improve quality" without evidence | Adds tokens/latency/cost past the task's need |

---

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · A pipeline routes everything to Opus 5. Cost is 4× budget; quality is acceptable. Which change best cuts cost while preserving quality on hard cases? (Select one)">
    A. Move all traffic to Haiku 4.5.
    B. Build a cascade: Haiku 4.5 / Sonnet 5 first, escalate to Opus 5 only when an output validation check fails, and cache the stable prefix.
    C. Escalate to Opus 5 whenever the model reports low confidence in its own answer.
    D. Increase effort to xhigh everywhere.

    **Answer: B.** Cheap-first with escalation on *validation* failure preserves quality on hard cases while most traffic runs cheaply; caching amortises the stable prefix. Blanket Haiku (A) sacrifices quality. Self-reported confidence (C) is an anti-pattern. Raising effort everywhere (D) increases cost.
  </AccordionItem>

  <AccordionItem title="Q2 · A team upgrades from Opus 4.6 to Opus 5 and their requests now return 400 errors. The requests set `budget_tokens` for thinking. What is the fix? (Select one)">
    A. Add more retries.
    B. Replace `budget_tokens` with `thinking: {\"type\": \"adaptive\"}` and control depth via `effort`, since `budget_tokens` is removed on Opus 5.
    C. Downgrade permanently to Haiku 4.5.
    D. Remove thinking entirely.

    **Answer: B.** `budget_tokens` is removed on Opus 5 (and Sonnet 5 / Fable 5.x) and returns 400; adaptive thinking plus `effort` is the supported replacement. Retries (A) won't fix a 400. Haiku (C) is the only model still using `budget_tokens` but is not an equivalent for Opus workloads. Removing thinking (D) discards needed reasoning.
  </AccordionItem>

  <AccordionItem title="Q3 · A refund agent must never issue refunds above $500. Where should this rule live? (Select one)">
    A. As a firmly worded sentence in the system prompt.
    B. As a programmatic tool-permission hook / validation that rejects any refund over \$500 before execution.
    C. As a few-shot example showing a refused large refund.
    D. In the model's thinking budget.

    **Answer: B.** Hard business rules require programmatic enforcement — a hook or validation the model cannot talk its way past. Prompt wording (A) and few-shot examples (C) are prompt-as-enforcement anti-patterns. Thinking budget (D) is unrelated.
  </AccordionItem>

  <AccordionItem title="Q4 · Prompt caching is enabled but hit rate is near zero. The prompt places the user's question first, then the system prompt and reference documents. Why, and what fixes it? (Select one)">
    A. Caching is broken; disable it.
    B. The volatile user turn sits before the stable content, so the cached prefix changes every call — reorder to stable-first (system → tools → docs) then the user turn, and mark the stable boundary with `cache_control`.
    C. The documents are too short.
    D. Haiku doesn't support caching.

    **Answer: B.** Caching keys on a stable prefix; putting the changing user turn first invalidates it every call. Stable-prefix-first with a `cache_control` boundary fixes it. Caching is not broken (A); document length (C) matters only for the ~1024-token minimum; Haiku does support caching (D) with a 2048-token minimum.
  </AccordionItem>

  <AccordionItem title="Q5 · Which are valid reasons NOT to stuff a whole 500k-token corpus into the 1M window every request? (Select two)">
    A. Higher token cost and latency per call.
    B. Possible accuracy degradation locating the relevant needle.
    C. The window physically cannot hold it.
    D. Structured outputs are disabled above 200k tokens.
    E. Caching is prohibited on large inputs.

    **Answer: A and B.** Stuffing raises cost and latency and can hurt retrieval accuracy within a huge context; retrieval (RAG) sends only relevant chunks. It fits the window (C is false at 500k of 1M), structured outputs are not size-gated that way (D), and caching is allowed on large inputs (E).
  </AccordionItem>

  <AccordionItem title="Q6 · In an active Fable 5.1 agentic session, the team wants to change the system instructions mid-run. What is the correct approach? (Select one)">
    A. Edit the original `system` field in place.
    B. Append the change as a new `role: \"system\"` message, leaving earlier turns untouched, because the harness must be append-only.
    C. Delete the earliest turns to make room.
    D. Reorder messages to put the new instruction first.

    **Answer: B.** Fable 5.1 thinking blocks are invalidated by editing/reordering/removing earlier turns, so mid-session changes are appended as a new system message. Editing in place (A), deleting turns (C) and reordering (D) all invalidate downstream thinking blocks.
  </AccordionItem>

  <AccordionItem title="Q7 · A team wants to force Fable 5.1 to always return a tool call using tool_choice set to any. It returns 400. What should they do? (Select one)">
    A. Retry until it works.
    B. Use `tool_choice: \"auto\"` with an instruction to use the tool, set `strict: true` on the tool schema, or use structured outputs — because forced tool use is unsupported on Fable 5.1.
    C. Switch to Haiku 4.5 permanently.
    D. Remove all tools.

    **Answer: B.** Forced tool use (`any` / forced tool) returns 400 on Fable 5.1; the supported paths are `auto` + instruction, `strict` schemas, or structured outputs. Retrying (A) won't fix a 400. Switching model (C) or removing tools (D) abandons the requirement.
  </AccordionItem>

  <AccordionItem title="Q8 · How should a new system-prompt version be rolled out to production? (Select one)">
    A. Swap it org-wide immediately to move fast.
    B. Validate offline against the golden set, canary on a small traffic slice with regression guardrails, ramp by percentage, and keep the prior version hot for instant rollback.
    C. Let each engineer edit the inline prompt in their own service.
    D. Ship it and monitor customer complaints.

    **Answer: B.** Prompts are governed, versioned assets: offline eval → canary → ramp → keep prior version for rollback. Org-wide swaps (A) and complaint-driven monitoring (D) skip validation. Per-engineer inline edits (C) destroy governance and reproducibility.
  </AccordionItem>

  <AccordionItem title="Q9 · A high-volume extraction subtask feeds a slower synthesis step. Which portfolio assignment is BEST? (Select one)">
    A. Opus 5 for both steps.
    B. Haiku 4.5 for the high-volume extraction; Sonnet 5 or Opus 5 for the synthesis — matching model cost to each step's difficulty.
    C. Fable 5.1 for both, for maximum quality.
    D. Haiku 4.5 for both, for maximum savings.

    **Answer: B.** Portfolio routing assigns the cheapest adequate model per step: Haiku for narrow high-volume extraction, a stronger model for harder synthesis. Opus/Fable for both (A, C) overpays; Haiku for both (D) risks the synthesis quality.
  </AccordionItem>

  <AccordionItem title="Q10 · Reusable capability blocks are being pasted into every prompt, bloating context and breaking the cache prefix. What is the better pattern? (Select one)">
    A. Package them as Skills (`SKILL.md`) loaded progressively on demand, keeping the stable cacheable prefix lean.
    B. Duplicate them into each service's prompt.
    C. Move them into the volatile user turn.
    D. Increase the context window.

    **Answer: A.** Skills load capability progressively on demand, keeping context lean and the cache prefix stable. Duplication (B) is what caused the bloat; moving them to the volatile turn (C) worsens caching; a bigger window (D) doesn't address cost or cache stability.
  </AccordionItem>

  <AccordionItem title="Q11 · Which statement about current-model thinking configuration is correct? (Select one)">
    A. All current models require `budget_tokens`.
    B. Current models use `thinking: {\"type\":\"adaptive\"}` with `effort` levels low/medium/high/xhigh; only Haiku 4.5 still uses `budget_tokens` and has no `effort` parameter.
    C. Effort only exists on Haiku 4.5.
    D. xhigh effort is the default on all models.

    **Answer: B.** Adaptive thinking plus effort is standard on current models; Haiku 4.5 is the exception still using `budget_tokens` and lacking `effort`. `budget_tokens` is not universal (A); effort is not Haiku-only (C); high (not xhigh) is the default (D).
  </AccordionItem>

  <AccordionItem title="Q12 · A prompt validated on Opus 5 is reused verbatim on Haiku 4.5 and quality drops. What is the correct lesson? (Select one)">
    A. Haiku 4.5 is defective.
    B. Prompts are validated per model version; a prompt tuned for one model must be re-tested (and often adjusted) against the golden set on any other model before use.
    C. Always use Opus 5.
    D. Quality drops are unavoidable and should be ignored.

    **Answer: B.** Model behaviour differs across the portfolio, so prompt versions carry the model they were validated against and must be re-evaluated when reused elsewhere. Haiku is not defective (A); mandating Opus (C) ignores cost; ignoring regressions (D) is negligent.
  </AccordionItem>

  <AccordionItem title="Q13 · A 4,000-token stable prefix on Sonnet 5 is reused with a 90% cache-hit rate at 10,000 req/day. Roughly what does caching save on input cost? (Select one)">
    A. Nothing; caching never helps large prefixes.
    B. Around 70%, because a cache read is ~0.1× input so the 4k prefix becomes ~400 effective tokens on hits.
    C. Exactly 50%, the Batch discount.
    D. 100%; cached requests are free.

    **Answer: B.** Read ≈ 0.1× input turns 4,000 prefix tokens into ~400 on the 90% of hits; the daily input cost drops from ~\$90 to ~\$27, about 70%. Caching does help large stable prefixes (A); 50% is the Batch discount, not caching (C); cached reads are cheap, not free (D).
  </AccordionItem>

  <AccordionItem title="Q14 · A pipeline must return a strictly typed JSON object for downstream systems, and it runs on Fable 5.1 where forcing tool use returns 400. What is the BEST approach? (Select one)">
    A. Force a tool call with `tool_choice: 'any'` and retry on 400.
    B. Use structured outputs with a `strict` JSON schema (or `auto` + instruction), which guarantees the shape without forcing tool use.
    C. Ask for JSON in the prompt and regex-parse the prose.
    D. Switch to free-text and reparse.

    **Answer: B.** Structured outputs with a `strict` schema conform by construction and don't require forced tool use, which Fable 5.1 rejects with 400. Forcing tools (A) fails; prompt-only JSON with regex (C) and free-text reparse (D) are brittle prose-parsing.
  </AccordionItem>

  <AccordionItem title="Q15 · A long agentic session on a thinking model overflows the window because tool results are huge. Which technique is correct, and what must be avoided? (Select one)">
    A. Delete the earliest user/assistant turns client-side.
    B. Use context editing to clear stale tool results server-side (and compaction for narrative), avoiding client-side edits to earlier reasoning that invalidate thinking blocks.
    C. Lower the temperature to shrink the context.
    D. Force the model to summarise itself in the same turn.

    **Answer: B.** Context editing removes stale tool results server-side; compaction summarises narrative — both avoid invalidating thinking blocks. Deleting turns client-side (A) breaks the append-only invariant; temperature (C) doesn't affect context size; in-turn self-summary (D) doesn't reclaim the window.
  </AccordionItem>

  <AccordionItem title="Q16 · A team runs every request on Opus 5 at `effort: xhigh` with acceptable quality and 5× budget; cost, not latency, is the complaint. Which TWO changes best cut cost while preserving quality? (Select two)">
    A. Cascade most traffic to Sonnet 5, escalating to Opus 5 only on a validation-check failure.
    B. Lower effort to an adequate level and re-run the per-segment regression suite.
    C. Escalate on the model's self-reported confidence.
    D. Move all traffic to Fable 5.1 for quality.
    E. Remove the eval suite to cut compute.

    **Answer: A and B.** A cheap-first cascade and right-sized effort (re-validated per segment) cut cost while preserving quality on hard cases. Self-report (C) is anti-pattern #4; Fable 5.1 (D) is the most expensive model; removing evals (E) removes the quality guard.
  </AccordionItem>

  <AccordionItem title="Q17 · Caching is enabled but the hit rate is ~3%. Investigation shows a per-request `request_id` string is prepended to the system prompt. What is the fix? (Select one)">
    A. Disable caching; it doesn't work here.
    B. Remove the volatile `request_id` from the prefix (log it separately) so the prefix is byte-stable, restoring cache hits.
    C. Shorten the documents.
    D. Increase the context window.

    **Answer: B.** A per-request token in the prefix changes it every call, so nothing caches; moving it out restores a stable prefix. Caching isn't broken (A); document length (C) only affects the minimum; a bigger window (D) is unrelated.
  </AccordionItem>

  <AccordionItem title="Q18 · A durable fact (a customer's contract tier) must persist across separate sessions without re-sending it in every prompt. Which mechanism fits, and what constraint applies? (Select one)">
    A. Paste the fact into every system prompt.
    B. Use the memory tool to persist the fact across sessions, but never store secrets/PII there inappropriately and keep it out of model-visible logs.
    C. Store it in `CLAUDE.local.md`.
    D. Increase retention to keep it in provider logs.

    **Answer: B.** The memory tool persists durable facts across sessions; the constraint is not to store secrets/PII inappropriately. Pasting per prompt (A) bloats context; `CLAUDE.local.md` (C) is a Claude Code dev file, not a runtime store; relying on provider retention (D) is not a memory mechanism and raises compliance risk.
  </AccordionItem>
</Accordions>

## Key takeaways

- Manage a **portfolio**: route/cascade to the cheapest model that clears the quality bar; escalate on **validation failure**, not self-reported confidence.
- Do the **cost math** — cascades and caching routinely cut spend 60–75% with quality preserved on hard cases.
- Manage **breaking changes**: `budget_tokens` removed (400) except on Haiku 4.5; forced tool use fails on Fable 5.1; re-run regressions before promoting a model.
- Treat prompts and guardrails as **versioned, governed assets**; enforce hard rules **programmatically**, never via prompt prose.
- Use **adaptive thinking + effort**; reserve `xhigh` for the hardest Opus/Fable work.
- **Cache stable-prefix-first**; keep reusable blocks as Skills; never place volatile content before the cached prefix.
- Fable 5.1 harnesses are **append-only**: freeze system/tools, append system messages, trim server-side, never force tools or silently downgrade.
- **Caching saving ∝ prefix size × hit rate**; a volatile token in the prefix (timestamps, request IDs) drops the hit rate to ~0 — remove it, don't disable caching.
- Guarantee output shape with **structured outputs / `strict` schemas**, especially on Fable 5.1 where forced tool use returns 400 — never parse prose.
- Manage long sessions with **context editing** (clear stale tool results) and **compaction** (summarise narrative) server-side; the **memory tool** persists durable facts (no secrets/PII).
