# D3 · Prompt Engineering and Structured Output

Prompts as system design, success criteria and rubrics, XML boundaries, few-shot and thinking, prefilling, structured output routes (output_config.format vs tool-use-as-schema vs strict tools), schema design, validation-retry, defensive parsing, extraction pipelines, and stop-reason handling.

import { Accordions, AccordionItem, Tabs, TabItem, Steps } from '@prosefly/astro-components';

This domain is worth roughly **12 of 60 items** and maps to the Structured Data Extraction and CI/CD scenarios. It treats a prompt as **system design**, not a text trick: stable structure for caching, explicit success criteria, clear content boundaries, and – above all – getting **reliable machine-readable output** out of the model and validating it. The most-tested decision is *which structured-output route to use*, including Fable 5.1's restriction that forbids forced tool choice.

## Learning objectives

By the end of this page you should be able to:

1. Treat a prompt as a **versioned, modular system** with a stable cacheable prefix and system/user role separation.
2. Write **explicit success criteria and rubrics** and use **XML tags** as content boundaries.
3. Select **few-shot examples** well and choose between **chain-of-thought and extended thinking**.
4. Use **prefilling** to steer output shape.
5. Choose the right **structured-output route** – `output_config.format` JSON schema, tool-use-as-schema, or strict tools – including the **Fable 5.1** no-forced-tool-choice rule.
6. Design **schemas** (enums, required, nullable, descriptions) and build a **validation-retry loop** with error feedback and **defensive parsing**.
7. Build **extraction pipelines** for heterogeneous documents with **per-type metrics** (anti-pattern 10) and handle **`refusal`** and **`max_tokens`** stop reasons.
8. Design **evaluator-optimizer** prompts and avoid **same-session self-review** (anti-pattern 9).

---

## 3.1 The prompt as system design

A production prompt is an engineered artefact with parts that change at different rates. Order them so the **stable** content sits first and can be cached.

```text
[ system role ]  stable: role, rules, output contract, tools  ← cache_control here
        +
[ user role ]    variable: the specific task, the document, the question
```

- **System vs user roles.** Put durable instructions, the output contract and constraints in the `system` prompt; put the specific task and data in `user` turns.
- **Stable prefix for caching.** Cache reads cost ~0.1× base input; put the system prompt, tool definitions and any long reference documents **first** and mark the last stable block with `cache_control: {"type": "ephemeral"}`. Minimum cacheable prefix ~1024 tokens (2048 on Haiku).
- **Modular assembly and versioning.** Build prompts from composable blocks (role, rules, schema, examples) and **version** them so you can evaluate and roll back. Treat prompt changes like code changes.

```python
system = [
    {"type": "text", "text": ROLE_AND_RULES},                       # stable
    {"type": "text", "text": OUTPUT_CONTRACT},                      # stable
    {"type": "text", "text": REFERENCE_DOC,
     "cache_control": {"type": "ephemeral"}},                       # stable, cache boundary
]
messages = [{"role": "user", "content": task_specific_input}]        # variable
```

:::caution[Fable 5.1 is append-only]
On Fable 5.1 you must **not** edit, reorder or remove earlier turns (it invalidates later thinking blocks). Freeze `system` and `tools`, put mid-session changes in `role: "system"` messages, and trim server-side via context editing/compaction. Design the stable prefix once and leave it alone. (Sonnet 5 does not allow mid-conversation system messages.)
:::

---

## 3.2 Explicit success criteria and rubrics

Vague instructions produce vague output. State **what good looks like** – measurable criteria and, for judged tasks, a rubric.

| Weak | Strong |
| --- | --- |
| "Summarise this well." | "Summarise in ≤150 words, lead with the decision, list exactly 3 risks, cite each figure's source." |
| "Check the code." | "Report every function lacking input validation as `{file, line, severity}`; severity ∈ high/medium/low." |

Rubrics also drive **evaluator-optimizer** loops (§3.10) and evals: the evaluator scores against the same explicit criteria the generator was given.

---

## 3.3 XML tags as content boundaries

Claude is trained to respect XML-style tags. Use them to separate instructions from data and to delimit output regions – this reduces injection risk and makes parsing deterministic.

```text
<instructions>Extract the invoice total. Return only the value.</instructions>
<document>
{ untrusted document text here }
</document>
```

Wrapping untrusted content in a clearly named tag both improves accuracy and signals to the model that the enclosed text is **data, not instructions** – a first-line defence against indirect prompt injection.

---

## 3.4 Few-shot example selection

Few-shot examples teach format and edge-case handling. Select them deliberately:

- **Representative** of the real distribution, including the hard/edge cases you care about.
- **Consistent** in format – the model mimics the shape, so any inconsistency propagates.
- **Diverse** enough to cover the classes, but not so many that they bloat context (and cost).
- **Correct** – a wrong example is worse than none.

For extraction, one example per document type usually beats ten examples of one type.

---

## 3.5 Chain-of-thought vs extended thinking

| Technique | What it is | Use when |
| --- | --- | --- |
| **Chain-of-thought (prompted)** | Ask the model to reason step by step in the output | You want visible reasoning you can inspect or when thinking is off |
| **Extended thinking** | Model reasons in dedicated thinking blocks before answering (`thinking: {"type": "adaptive"}`) | Hard reasoning, multi-step problems, agentic planning |

On current models thinking is `adaptive`; `budget_tokens` is removed on Fable 5.x / Opus 5 / Sonnet 5 (400 error) and only **Haiku 4.5** still uses `budget_tokens`. Effort levels are `low | medium | high (default) | xhigh` – use `xhigh` for the hardest coding/agentic work on Opus 5 / Fable 5.1. On Fable 5.1 thinking is always on.

:::caution[Thinking blocks and fallback]
Thinking blocks are readable only by the producing model or a newer one. If you fall back from Fable 5.1 to an older model, the earlier thinking blocks are silently dropped – design for this in Domain 5's fallback discussion.
:::

---

## 3.6 Prefilling

Prefill the start of the assistant turn to constrain the shape of the output – force JSON, skip preamble, or lock a format.

```python
messages = [
    {"role": "user", "content": "Return the extracted fields."},
    {"role": "assistant", "content": "{"},        # prefill: forces the model to continue JSON
]
```

Prefilling is a lightweight steer, not a guarantee. For strong guarantees use structured outputs or strict tools (next). Note: prefilling interacts poorly with extended thinking on models where thinking must come first – prefer structured outputs there.

---

## 3.7 Structured output: the three routes

This is the highest-value decision table in the domain.

| Route | How | Guarantee | Best for | Caveats |
| --- | --- | --- | --- | --- |
| **`output_config.format` (structured outputs)** | Pass a JSON schema in `output_config.format` | Model output conforms to the schema | You want a typed response object, no tool semantics | Newer capability; check model support |
| **Tool-use-as-schema** | Define a tool whose `input_schema` is your target shape; read the tool call's input | Strong when combined with `strict: true` | You already use tools, or want the model to "emit" a record | Historically paired with forced `tool_choice` – restricted on Fable 5.1 |
| **Strict tools (`strict: true`)** | Add `strict: true` to a tool schema | Enforces the schema on the tool input | Reliable tool arguments / record emission | Preferred on Fable 5.1 where forcing tool_choice is blocked |

```python
# Route 1: structured outputs via output_config.format
resp = client.messages.create(
    model="claude-opus-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": f"<document>{doc}</document> Extract fields."}],
    output_config={"format": {"type": "json_schema", "schema": INVOICE_SCHEMA}},
)
```

```python
# Route 3: strict tool as a schema (works on Fable 5.1 with tool_choice: auto)
tools = [{
    "name": "record_invoice",
    "description": "Record the extracted invoice fields.",
    "strict": True,
    "input_schema": INVOICE_SCHEMA,
}]
resp = client.messages.create(
    model="claude-fable-5-1",
    max_tokens=1024,
    tools=tools,
    tool_choice={"type": "auto"},   # NOT {"type":"tool"} or "any" on Fable 5.1
    messages=[{"role": "user", "content": prompt}],
)
```

:::danger[Fable 5.1 · no forced tool choice]
On **Claude Fable 5.1**, `tool_choice: "any"` and forced `{"type": "tool", "name": …}` return **400**. To get structured output on Fable 5.1, use one of: `tool_choice: "auto"` **plus an instruction** to call the tool, `strict: true` tool schemas, or **structured outputs** (`output_config.format`). This is a favourite distractor: any option that forces a tool on Fable 5.1 is wrong.
:::

:::tip[Exam signal]
"Guaranteed to match a schema, no tool semantics needed" → **structured outputs**. "Already using tools / want the model to emit a record" → **tool-use-as-schema with `strict: true`**. "On Fable 5.1" → never force `tool_choice`; use auto + instruction, strict, or structured outputs.
:::

---

## 3.8 Schema design

A good schema is self-documenting and constraining.

```json
{
  "type": "object",
  "properties": {
    "invoice_number": { "type": "string", "description": "The vendor invoice ID as printed." },
    "total": { "type": "number", "description": "Grand total in the invoice currency." },
    "currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] },
    "due_date": { "type": ["string", "null"], "description": "ISO 8601 date, or null if absent." },
    "line_items": { "type": "array", "items": { "type": "object" } }
  },
  "required": ["invoice_number", "total", "currency"],
  "additionalProperties": false
}
```

- **`enum`** constrains a field to known values (prevents free-text drift).
- **`required`** forces presence; omit optional fields from it.
- **Nullable** via `"type": ["string", "null"]` – model the "absent" case explicitly rather than inviting a hallucinated value.
- **`description`** on every field – the model reads these; they are the cheapest accuracy lever.
- **`additionalProperties: false`** keeps output tight.

---

## 3.9 Validation-retry loop and defensive parsing

Even with schemas, validate and be ready to retry with **specific error feedback**.

```python
import json, jsonschema

def extract_with_retry(client, prompt, schema, max_attempts=3):
    messages = [{"role": "user", "content": prompt}]
    for attempt in range(max_attempts):
        resp = client.messages.create(
            model="claude-opus-5", max_tokens=1024, messages=messages,
            output_config={"format": {"type": "json_schema", "schema": schema}},
        )
        if resp.stop_reason == "refusal":
            return {"status": "refused"}                 # do not retry blindly
        if resp.stop_reason == "max_tokens":
            raise OutputTruncated()                      # raise max_tokens and retry
        text = "".join(b.text for b in resp.content if b.type == "text")
        try:
            data = json.loads(text)                      # defensive parse
            jsonschema.validate(data, schema)
            return {"status": "ok", "data": data}
        except (json.JSONDecodeError, jsonschema.ValidationError) as e:
            messages.append({"role": "assistant", "content": text})
            messages.append({"role": "user",
                             "content": f"That output failed validation: {e}. "
                                        f"Return only valid JSON matching the schema."})
    return {"status": "failed_validation"}
```

Key points: **feed the specific error back** so the model self-corrects; parse defensively (never `eval`); treat `refusal` and `max_tokens` as distinct outcomes, not validation failures.

---

## 3.10 Extraction pipelines for heterogeneous documents

When you extract from many document *types* (invoices, contracts, receipts), a single aggregate accuracy number **hides** a type that is failing badly.

:::danger[Anti-pattern 10 · Aggregate metrics masking per-type failure]
Reporting "94% overall extraction accuracy" can hide that contracts are at 60% while invoices are at 99%. Always measure **per document type** (per-segment metrics). The exam-correct evaluation slices the metric by type and gates on the worst type.
:::

```text
                 ┌─► classify type ─┬─► invoice schema  ─► validate ─► metrics[invoice]
Document ─► route ┤                 ├─► contract schema ─► validate ─► metrics[contract]
                 └─────────────────┴─► receipt schema  ─► validate ─► metrics[receipt]
                          per-type accuracy, not one aggregate number
```

Design: route by type → apply the type-specific schema → validate → record per-type metrics. Report and alert per type.

---

## 3.11 Handling refusal and max_tokens

| stop_reason | Meaning | Correct handling |
| --- | --- | --- |
| `refusal` | Model declined (safety) | Stop; do not "retry harder"; escalate or reframe legitimately; log |
| `max_tokens` | Output truncated | Not a completion; raise `max_tokens`, or chunk the task, then retry |
| `end_turn` | Model finished | Proceed |
| `tool_use` | Wants a tool | Run tool, continue loop |

Treating `max_tokens` output as a complete, parseable result is a silent-failure trap.

---

## 3.12 Evaluator-optimizer prompts and independence

For a generate-critique-refine loop, give the evaluator the **explicit rubric** and run it as an **independent context** – a fresh session, ideally a different model. Reusing the generating conversation to grade itself keeps the reasoning bias that produced the flaw.

:::danger[Anti-pattern 9 · Same-session self-review]
"Ask the model in the same chat whether its answer is correct" retains reasoning-context bias. Use a separate session / different model as the judge, scoring against the written rubric. This applies to LLM-as-judge evals too.
:::

---

## 3.13 Schema design patterns for hard extraction cases

Beyond enums and nullable types, several schema patterns decide whether extraction is reliable on messy documents.

| Pattern | JSON Schema | Solves |
| --- | --- | --- |
| **Discriminated union** | `oneOf` with a `const` discriminator field | One tool/endpoint handling several document types cleanly |
| **Bounded arrays** | `"maxItems": N` on `line_items` | Runaway output and truncation on huge tables |
| **Formatted strings** | `"format": "date"`, `"pattern": "^[A-Z]{3}$"` | Dates/codes that must match a shape |
| **Confidence & provenance** | `confidence` enum + `source_span` field | Downstream gating and human verification |
| **Explicit "not found"** | nullable + a `found: boolean` sibling | Distinguishing "absent" from "missed" |

```json
{
  "type": "object",
  "oneOf": [
    { "properties": { "doc_type": { "const": "invoice" }, "total": { "type": "number" } },
      "required": ["doc_type", "total"] },
    { "properties": { "doc_type": { "const": "contract" }, "term_months": { "type": "integer" } },
      "required": ["doc_type", "term_months"] }
  ]
}
```

:::tip[Exam signal]
"The optional field is sometimes hallucinated when absent" → model it as **nullable with an explicit found flag**, not as a plain optional the model feels pressure to fill. "One pipeline, several document types" → **discriminated union**, not one loose object.
:::

---

## 3.14 A robust validation-retry harness (with schema-error feedback)

The exam-correct loop treats `refusal` and `max_tokens` as distinct outcomes, feeds the *specific* validation error back, and never parses unsafely.

<Tabs>
  <TabItem label="Python">

```python
import json, jsonschema
from jsonschema import Draft202012Validator

def extract(client, prompt, schema, model="claude-opus-5", max_attempts=3):
    messages = [{"role": "user", "content": prompt}]
    for attempt in range(max_attempts):
        resp = client.messages.create(
            model=model, max_tokens=2048, messages=messages,
            output_config={"format": {"type": "json_schema", "schema": schema}},
        )
        if resp.stop_reason == "refusal":
            return {"status": "refused"}                       # safety stop — do not bypass
        if resp.stop_reason == "max_tokens":
            return {"status": "truncated", "retry": "raise_limit_or_chunk"}  # not complete
        text = "".join(b.text for b in resp.content if b.type == "text")
        try:
            data = json.loads(text)                             # never eval()
            errors = sorted(Draft202012Validator(schema).iter_errors(data),
                            key=lambda e: e.path)
            if not errors:
                return {"status": "ok", "data": data}
            detail = "; ".join(f"{list(e.path)}: {e.message}" for e in errors[:5])
        except json.JSONDecodeError as e:
            detail = f"invalid JSON at pos {e.pos}: {e.msg}"
        messages += [
            {"role": "assistant", "content": text},
            {"role": "user", "content": f"Validation failed: {detail}. "
                                        f"Return ONLY valid JSON matching the schema."},
        ]
    return {"status": "failed_validation"}
```

  </TabItem>
  <TabItem label="TypeScript">

```typescript
import Ajv from 'ajv';
const ajv = new Ajv({ allErrors: true });

async function extract(client, prompt, schema, model = 'claude-opus-5', maxAttempts = 3) {
  const validate = ajv.compile(schema);
  const messages: any[] = [{ role: 'user', content: prompt }];
  for (let i = 0; i < maxAttempts; i++) {
    const resp = await client.messages.create({
      model, max_tokens: 2048, messages,
      output_config: { format: { type: 'json_schema', schema } },
    });
    if (resp.stop_reason === 'refusal') return { status: 'refused' };
    if (resp.stop_reason === 'max_tokens') return { status: 'truncated' }; // not complete
    const text = resp.content.filter(b => b.type === 'text').map(b => b.text).join('');
    try {
      const data = JSON.parse(text);                    // never eval / Function
      if (validate(data)) return { status: 'ok', data };
      const detail = (validate.errors ?? []).slice(0, 5)
        .map(e => `${e.instancePath} ${e.message}`).join('; ');
      messages.push({ role: 'assistant', content: text },
        { role: 'user', content: `Validation failed: ${detail}. Return ONLY valid JSON.` });
    } catch (e) {
      messages.push({ role: 'assistant', content: text },
        { role: 'user', content: `Invalid JSON: ${e}. Return ONLY valid JSON.` });
    }
  }
  return { status: 'failed_validation' };
}
```

  </TabItem>
</Tabs>

:::danger[Never `eval` model output]
Parsing model output with `eval()` (Python) or `Function`/`eval` (JS) executes arbitrary code from an untrusted source. Always `json.loads` / `JSON.parse` behind a schema validator.
:::

---

## 3.15 Structured-output cost and token arithmetic

Structured output is not free of token cost, and schema verbosity shows up on the bill. A quick model:

```text
Prompt = system(1,500) + tools/schema(800) + document(6,000) = 8,300 input tokens
Output = ~400 tokens of JSON

On Opus 5 ($5/M in, $25/M out):
  input  = 8,300 × $5  / 1e6 = $0.0415
  output =   400 × $25 / 1e6 = $0.0100
  per doc                     ≈ $0.0515

Cache the stable 2,300-token prefix (system + schema), read ≈ 0.1×:
  cached prefix = 2,300 × $0.5 / 1e6 = $0.00115 (vs $0.0115 uncached)
  → saves ≈ $0.0104 per doc; over 1M docs ≈ $10,400 saved.

Batch API (latency-tolerant) halves the remaining cost → ~$0.026/doc.
```

| Lever | Effect | When correct |
| --- | --- | --- |
| Cache the schema + system prefix | ~10× cheaper reads on the stable part | Repeated extraction with a fixed schema |
| Batch API | 50% off, ≤24h | Latency-tolerant bulk extraction |
| Cheaper model (Haiku 4.5) | Lower per-token price | Simple, well-constrained schemas |
| Tighter schema (`maxItems`, closed objects) | Fewer output tokens, less truncation | Large tables/line items |

:::tip[Exam signal]
"MOST cost-effective way to extract from a million documents overnight" combines **caching the stable schema/system prefix** with the **Batch API**, and possibly a cheaper model — not "use a bigger model" or "call in real time at high concurrency".
:::

---

## Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "Prefilling guarantees valid JSON." | Prefill only *steers*; structured outputs / strict tools *guarantee* the schema. | Route-selection items. |
| "Force `tool_choice` for reliable tool output." | On Fable 5.1 forcing tool choice is a 400; use auto+instruction, strict, or structured outputs. | The signature Fable 5.1 distractor. |
| "`max_tokens` output is just a bit shorter." | It is *truncated*; parsing it as complete corrupts data. | Silent-truncation trap. |
| "A refusal is a validation error to retry." | Refusal is a safety stop; retrying to bypass is wrong. | `stop_reason` handling items. |
| "One aggregate accuracy number is enough." | Aggregate hides a failing document type; slice per type (#10). | Evaluation items. |
| "The model can grade its own answer in-session." | Same-session self-review inherits the bias (#9); use an independent judge. | Evaluator-optimizer / LLM-as-judge items. |
| "`budget_tokens` controls thinking everywhere." | Only Haiku 4.5 uses `budget_tokens`; Opus/Sonnet/Fable use adaptive thinking + `effort`. | Model-facts distractor. |
| "Variable input first is fine." | The stable prefix must be first to be cacheable. | Caching items. |

---

## Scenario walkthrough — an extraction pipeline that looks fine but ships bad contracts

**Situation.** A pipeline extracts fields from invoices, contracts and receipts on Fable 5.1 and writes to a database. It reports **96% overall accuracy**, yet the legal team keeps finding wrong contract term lengths. The engineers force `tool_choice` onto a `record` tool and see intermittent 400s; when a big contract is processed, the JSON sometimes ends mid-array and the code writes the partial object. A proposal is to "grade outputs by asking the model, in the same chat, if it is confident."

**Expert reasoning trace.**

<Steps>

1. **Fix the 400s first.** Fable 5.1 forbids forced `tool_choice`. Switch to **structured outputs (`output_config.format`)** or a **`strict: true` tool with `tool_choice: "auto"` plus an instruction**. Reject "retry with backoff" (a 400 is deterministic) and `"any"` (also blocked).

2. **Fix the silent truncation.** JSON ending mid-array with `stop_reason == max_tokens` is **truncation**, not a result. Raise `max_tokens`, or **chunk** the contract (or bound `line_items` with `maxItems`), then retry. Never write the partial object.

3. **Fix the metric.** 96% overall **masks** a failing type (#10). Report **per-document-type accuracy**, alert on contracts, and **gate release on the worst type**. Reject "bigger sample" and "average more runs" — both still aggregate.

4. **Fix the evaluation design.** Same-session self-confidence is **#9** and also **self-reported confidence** territory. Use an **independent judge** (fresh session, ideally a different model) scoring against a written rubric.

5. **Harden the schema.** Contracts need a **discriminated union** branch with `term_months` as a typed integer and a **nullable + found flag** for optional clauses so absence is not hallucinated.

</Steps>

**Exam-correct decision:** structured outputs / strict tools (not forced choice), truncation handling by raising/chunking, per-type metrics gating on the worst type, an independent evaluator, and a discriminated-union schema. Each rejected option is a named trap (Fable forced choice, silent truncation, aggregate masking #10, same-session #9).

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| Force `tool_choice: "any"`/`{"type":"tool"}` on Fable 5.1 | Returns 400; use auto+instruction, strict, or structured outputs |
| Put the variable task before the stable system prompt | Breaks caching; stable content must be first |
| Use one aggregate accuracy metric across document types | Hides per-type failure (anti-pattern 10) |
| Treat `max_tokens` output as a complete result | It is truncated; raise the limit or chunk |
| Retry a `refusal` by rephrasing to bypass safety | Refusals are safety stops, not validation errors |
| Grade an answer in the same session that produced it | Same-session self-review bias (anti-pattern 9) |
| Set `budget_tokens` on Opus 5 / Sonnet 5 / Fable 5.1 | Removed (400); only Haiku 4.5 uses it |
| Rely on prefilling as a hard guarantee | It steers, not guarantees; use structured outputs/strict tools |
| Skip field descriptions in the schema | Descriptions are the cheapest accuracy lever |
| Edit earlier turns mid-session on Fable 5.1 | Invalidates thinking blocks; harness must be append-only |
| Parse model output with `eval()` | Executes untrusted code; use `json.loads`/`JSON.parse` behind a validator |
| Model an optional field as a plain optional | The model may hallucinate a value; use nullable + a found flag |
| Use one loose object for several document types | Use a discriminated union (`oneOf` + `const`) |
| Leave line-item arrays unbounded | Add `maxItems` to prevent runaway output and truncation |
| Real-time high concurrency for bulk overnight extraction | Use the Batch API (50% off, ≤24h) and cache the schema prefix |

---

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · A pipeline on Claude Fable 5.1 must return records that match a JSON schema. An engineer sets a forced tool_choice on the record tool and gets 400 errors. What is the correct fix? (Select one)">
    A. Retry with exponential backoff.
    B. Use `tool_choice: 'auto'` with an instruction to call the tool, or `strict: true` tools, or structured outputs via `output_config.format`.
    C. Switch to `tool_choice: 'any'`.
    D. Lower max_tokens.

    **Answer: B.** Fable 5.1 rejects forced tool choice (`any` and `{type:"tool"}` both 400). The supported routes are auto+instruction, strict tools, or structured outputs. Backoff (A) does not fix a 400; `any` (C) is also blocked; max_tokens (D) is unrelated.
  </AccordionItem>

  <AccordionItem title="Q2 · A team reports 94% overall extraction accuracy across invoices, contracts and receipts. A stakeholder complains contracts are frequently wrong. What is the BEST evaluation change? (Select one)">
    A. Increase the overall sample size.
    B. Report per-document-type accuracy and gate on the worst-performing type.
    C. Raise the model temperature for contracts.
    D. Average three runs to smooth the metric.

    **Answer: B.** Aggregate accuracy masks a failing type (anti-pattern 10). Per-type metrics expose and gate on the weak type. Larger samples (A) still aggregate; temperature (C) does not fix accuracy; averaging (D) further hides the problem.
  </AccordionItem>

  <AccordionItem title="Q3 · A prompt places the specific user document first and the long stable system rules last. Latency and cost are high due to no cache hits. What is the fix? (Select one)">
    A. Shorten the document.
    B. Put the stable system prompt, tools and reference material first with `cache_control` on the last stable block, and the variable task after.
    C. Disable caching.
    D. Use a bigger model.

    **Answer: B.** Caching requires the stable content first; the variable task goes last. Shortening the document (A) does not enable caching, disabling caching (C) worsens cost, and model size (D) is irrelevant.
  </AccordionItem>

  <AccordionItem title="Q4 · An extraction returns text ending mid-object and `stop_reason` is `max_tokens`. The code JSON-parses it and fails. What is the correct handling? (Select one)">
    A. Treat the partial JSON as the result.
    B. Recognise `max_tokens` as truncation, raise the token limit or chunk the task, then retry — do not parse it as complete.
    C. Log a generic error and return empty.
    D. Ask the model in the same chat if it is sure.

    **Answer: B.** `max_tokens` means truncated output, not completion. Raise the limit or chunk. Parsing partial output (A) is a silent-failure trap; returning empty (C) is anti-pattern 7; same-session self-check (D) is unrelated.
  </AccordionItem>

  <AccordionItem title="Q5 · Which TWO schema design choices most improve extraction reliability? (Select two)">
    A. A `description` on every field.
    B. Free-text for fields that have a fixed set of values.
    C. An `enum` for the currency field and explicit nullable types for optional fields.
    D. Omitting `required` entirely.
    E. Allowing `additionalProperties: true`.

    **Answer: A and C.** Descriptions guide the model and enums/nullable types constrain and model absence explicitly. Free text (B) invites drift, omitting `required` (D) loses presence guarantees, and open additional properties (E) loosens the output.
  </AccordionItem>

  <AccordionItem title="Q6 · A validation-retry loop keeps failing. Currently it just re-sends the same prompt. What change most helps the model self-correct? (Select one)">
    A. Increase the number of retries only.
    B. Feed the specific validation error back to the model and ask it to return only valid JSON matching the schema.
    C. Switch to `eval()` to parse the output.
    D. Lower max_tokens.

    **Answer: B.** Specific error feedback lets the model fix the exact problem. More blind retries (A) rarely help, `eval()` (C) is unsafe, and lowering max_tokens (D) risks truncation.
  </AccordionItem>

  <AccordionItem title="Q7 · When should you use extended thinking rather than prompted chain-of-thought? (Select one)">
    A. For every request, always.
    B. For hard multi-step reasoning or agentic planning, using `thinking: {'type': 'adaptive'}` on current models.
    C. Never; chain-of-thought is always better.
    D. Only to reduce cost.

    **Answer: B.** Extended thinking suits hard, multi-step or agentic reasoning. It is not needed for every request (A), is not universally worse (C), and does not reduce cost (D) — thinking tokens add cost.
  </AccordionItem>

  <AccordionItem title="Q8 · An evaluator-optimizer loop grades the translation in the same conversation that produced it and quality does not improve. What is the flaw? (Select one)">
    A. The rubric is too detailed.
    B. Same-session self-review retains reasoning bias; run the evaluator as an independent context, ideally a different model, against the explicit rubric.
    C. The generator needs a higher temperature.
    D. There are too few retries.

    **Answer: B.** Anti-pattern 9. Independence removes the shared bias. A detailed rubric (A) is good, temperature (C) does not fix bias, and more retries (D) repeat the biased judge.
  </AccordionItem>

  <AccordionItem title="Q9 · A document contains untrusted third-party text that instructs the model to ignore its task. What prompt-design practice reduces this risk? (Select one)">
    A. Concatenate the document directly into the instructions.
    B. Wrap the untrusted content in a clearly named XML tag (e.g. `<document>…</document>`) and instruct the model to treat it as data, not instructions.
    C. Trust the model to notice.
    D. Set temperature to 0.

    **Answer: B.** XML content boundaries separate data from instructions and blunt indirect injection. Concatenation (A) invites injection, trusting the model (C) is not a control, and temperature (D) is irrelevant.
  </AccordionItem>

  <AccordionItem title="Q10 · On Claude Sonnet 5, a harness tries to inject a mid-conversation system message. What is true? (Select one)">
    A. Sonnet 5 allows mid-conversation system messages.
    B. Sonnet 5 does not allow mid-conversation system messages; design the system prompt up front.
    C. Only Fable 5.1 forbids this.
    D. Set `budget_tokens` to enable it.

    **Answer: B.** Sonnet 5 does not allow mid-conversation system messages, so the design must fix the system prompt up front. A is false; Fable 5.1's append-only constraint is separate (C); `budget_tokens` (D) is removed on Sonnet 5.
  </AccordionItem>

  <AccordionItem title="Q11 · A pipeline needs a guaranteed schema-conformant object and does not use tools for anything else. Which route is cleanest? (Select one)">
    A. Prefill the assistant turn with `{` and hope.
    B. Structured outputs via `output_config.format` with a JSON schema.
    C. Force `tool_choice` to a dummy tool on Fable 5.1.
    D. Parse free-text prose with a regex.

    **Answer: B.** Structured outputs give a schema guarantee without tool semantics. Prefill (A) only steers, forcing tool_choice on Fable 5.1 (C) 400s, and regex on prose (D) is brittle.
  </AccordionItem>

  <AccordionItem title="Q13 · One extraction pipeline must handle invoices, contracts and receipts, each with different required fields. Which schema pattern is BEST? (Select one)">
    A. One loose object with every possible field optional.
    B. A discriminated union (`oneOf` with a `const` `doc_type` discriminator), so each branch enforces its own required fields.
    C. Three unrelated endpoints with no shared contract.
    D. A single string field holding raw JSON text.

    **Answer: B.** A discriminated union cleanly enforces per-type requirements in one schema. A loosens everything and invites drift; C loses a shared contract; D abandons schema guarantees entirely.
  </AccordionItem>

  <AccordionItem title="Q14 · An optional 'renewal_clause' field is frequently hallucinated when the clause is absent. Which schema change MOST reduces this? (Select one)">
    A. Make it a required string so the model always fills it.
    B. Model it as nullable (`[\"string\", \"null\"]`) with an explicit `renewal_found: boolean` sibling, so absence is represented rather than invented.
    C. Remove the field entirely.
    D. Raise the temperature so answers vary.

    **Answer: B.** Nullable plus an explicit found flag lets the model represent absence instead of hallucinating. Required (A) forces a value; removing it (C) loses the data; temperature (D) worsens consistency.
  </AccordionItem>

  <AccordionItem title="Q15 · A million documents must be extracted overnight as cheaply as possible with a fixed schema. Which combination is MOST cost-effective? (Select two)">
    A. Cache the stable system + schema prefix so reads cost ~0.1x.
    B. Use the Message Batches API for the latency-tolerant bulk run (50% off, ≤24h).
    C. Call the real-time API at maximum concurrency.
    D. Use Fable 5.1 at $10/$50 for every document.
    E. Randomise the prompt each call to avoid stale reads.

    **Answer: A and B.** Caching the fixed prefix and the Batch API together cut cost sharply for latency-tolerant bulk work. Real-time high concurrency (C) costs full price and risks limits; the most expensive model (D) raises cost; randomising (E) destroys the cacheable prefix.
  </AccordionItem>

  <AccordionItem title="Q16 · A developer parses extraction output with `eval()` because 'it handles trailing commas'. What is the correct critique? (Select one)">
    A. It is fine if the model is trusted.
    B. `eval()` executes arbitrary untrusted code from the model; parse with `json.loads`/`JSON.parse` behind a schema validator and feed validation errors back on failure.
    C. Use a regex instead of `eval()`.
    D. Lower max_tokens so the output is smaller.

    **Answer: B.** Model output is untrusted; `eval()` is a code-execution risk. Safe parsing plus schema validation and error feedback is correct. The model is never 'trusted' for eval (A); a regex (C) is brittle; max_tokens (D) is unrelated.
  </AccordionItem>

  <AccordionItem title="Q17 · A large contract's extracted JSON ends mid-array with `stop_reason` `max_tokens`, and the pipeline writes the partial object. What is the correct handling? (Select one)">
    A. Write the partial object; it is mostly complete.
    B. Treat `max_tokens` as truncation: raise the limit or chunk the document (or bound arrays with `maxItems`), then retry — never persist the partial output.
    C. Return an empty object so the pipeline continues.
    D. Ask the model in the same session whether it finished.

    **Answer: B.** `max_tokens` is truncation, not completion. Persisting the partial (A) corrupts data; returning empty (C) is silent failure (#7); same-session self-check (D) does not address truncation.
  </AccordionItem>

  <AccordionItem title="Q18 · On Opus 5 an agentic harness sets `budget_tokens` for thinking and receives a 400. What is true and what is the fix? (Select one)">
    A. `budget_tokens` works on Opus 5; retry the 400.
    B. `budget_tokens` is Haiku-4.5-only; on Opus 5 use adaptive thinking with an `effort` level (`low|medium|high|xhigh`), e.g. `xhigh` for the hardest work.
    C. Disable thinking to avoid the error.
    D. Switch to `tool_choice: any` to enable budgets.

    **Answer: B.** Only Haiku 4.5 uses `budget_tokens`; Opus 5 uses effort levels. The 400 is deterministic, not transient (A); disabling thinking (C) removes needed reasoning; tool_choice (D) is unrelated to thinking control.
  </AccordionItem>
</Accordions>

## Key takeaways

- Design prompts as versioned, modular systems: stable content first (system role, tools, docs) with `cache_control`; variable task last.
- State explicit success criteria and rubrics; use XML tags to separate untrusted data from instructions.
- Pick structured output deliberately: `output_config.format` for schema guarantees, tool-use-as-schema/`strict` for record emission; never force `tool_choice` on Fable 5.1.
- Design schemas with enums, required, nullable types, descriptions, discriminated unions and bounded arrays; model absence with nullable + a found flag.
- Build a validation-retry loop that feeds the *specific* schema error back and parses safely (never `eval`); treat `refusal` and `max_tokens` as distinct outcomes.
- Slice extraction metrics per document type and gate on the worst; never trust one aggregate number.
- Control reasoning with adaptive thinking + `effort` (Opus/Sonnet/Fable); only Haiku 4.5 uses `budget_tokens`.
- Run evaluators as independent contexts; on Fable 5.1 keep the harness append-only and freeze system/tools.
- For bulk extraction, cache the stable schema/system prefix and use the Batch API (50% off) — the cost levers, not a bigger model.
