AI Cert Prep
Type to search documentation.

Domains

D3 · Prompt Engineering and Structured Output

Prompts as system design, success criteria and rubrics, XML boundaries, few-shot and thinking, prefilling, structured output routes (output_config.format vs tool-use-as-schema vs strict tools), schema design, validation-retry, defensive parsing, extraction pipelines, and stop-reason handling.

This domain is worth roughly 12 of 60 items and maps to the Structured Data Extraction and CI/CD scenarios. It treats a prompt as system design, not a text trick: stable structure for caching, explicit success criteria, clear content boundaries, and – above all – getting reliable machine-readable output out of the model and validating it. The most-tested decision is which structured-output route to use, including Fable 5.1’s restriction that forbids forced tool choice.

Learning objectives

By the end of this page you should be able to:

  1. Treat a prompt as a versioned, modular system with a stable cacheable prefix and system/user role separation.
  2. Write explicit success criteria and rubrics and use XML tags as content boundaries.
  3. Select few-shot examples well and choose between chain-of-thought and extended thinking.
  4. Use prefilling to steer output shape.
  5. Choose the right structured-output route – output_config.format JSON schema, tool-use-as-schema, or strict tools – including the Fable 5.1 no-forced-tool-choice rule.
  6. Design schemas (enums, required, nullable, descriptions) and build a validation-retry loop with error feedback and defensive parsing.
  7. Build extraction pipelines for heterogeneous documents with per-type metrics (anti-pattern 10) and handle refusal and max_tokens stop reasons.
  8. Design evaluator-optimizer prompts and avoid same-session self-review (anti-pattern 9).

3.1 The prompt as system design

A production prompt is an engineered artefact with parts that change at different rates. Order them so the stable content sits first and can be cached.

text
[ system role ] stable: role, rules, output contract, tools ← cache_control here
+
[ user role ] variable: the specific task, the document, the question
  • System vs user roles. Put durable instructions, the output contract and constraints in the system prompt; put the specific task and data in user turns.
  • Stable prefix for caching. Cache reads cost ~0.1× base input; put the system prompt, tool definitions and any long reference documents first and mark the last stable block with cache_control: {"type": "ephemeral"}. Minimum cacheable prefix ~1024 tokens (2048 on Haiku).
  • Modular assembly and versioning. Build prompts from composable blocks (role, rules, schema, examples) and version them so you can evaluate and roll back. Treat prompt changes like code changes.
python
system = [
{"type": "text", "text": ROLE_AND_RULES}, # stable
{"type": "text", "text": OUTPUT_CONTRACT}, # stable
{"type": "text", "text": REFERENCE_DOC,
"cache_control": {"type": "ephemeral"}}, # stable, cache boundary
]
messages = [{"role": "user", "content": task_specific_input}] # variable

Fable 5.1 is append-only

On Fable 5.1 you must not edit, reorder or remove earlier turns (it invalidates later thinking blocks). Freeze system and tools, put mid-session changes in role: "system" messages, and trim server-side via context editing/compaction. Design the stable prefix once and leave it alone. (Sonnet 5 does not allow mid-conversation system messages.)


3.2 Explicit success criteria and rubrics

Vague instructions produce vague output. State what good looks like – measurable criteria and, for judged tasks, a rubric.

WeakStrong
“Summarise this well.”“Summarise in ≤150 words, lead with the decision, list exactly 3 risks, cite each figure’s source.”
“Check the code.”“Report every function lacking input validation as {file, line, severity}; severity ∈ high/medium/low.”

Rubrics also drive evaluator-optimizer loops (§3.10) and evals: the evaluator scores against the same explicit criteria the generator was given.


3.3 XML tags as content boundaries

Claude is trained to respect XML-style tags. Use them to separate instructions from data and to delimit output regions – this reduces injection risk and makes parsing deterministic.

text
<instructions>Extract the invoice total. Return only the value.</instructions>
<document>
{ untrusted document text here }
</document>

Wrapping untrusted content in a clearly named tag both improves accuracy and signals to the model that the enclosed text is data, not instructions – a first-line defence against indirect prompt injection.


3.4 Few-shot example selection

Few-shot examples teach format and edge-case handling. Select them deliberately:

  • Representative of the real distribution, including the hard/edge cases you care about.
  • Consistent in format – the model mimics the shape, so any inconsistency propagates.
  • Diverse enough to cover the classes, but not so many that they bloat context (and cost).
  • Correct – a wrong example is worse than none.

For extraction, one example per document type usually beats ten examples of one type.


3.5 Chain-of-thought vs extended thinking

TechniqueWhat it isUse when
Chain-of-thought (prompted)Ask the model to reason step by step in the outputYou want visible reasoning you can inspect or when thinking is off
Extended thinkingModel reasons in dedicated thinking blocks before answering (thinking: {"type": "adaptive"})Hard reasoning, multi-step problems, agentic planning

On current models thinking is adaptive; budget_tokens is removed on Fable 5.x / Opus 5 / Sonnet 5 (400 error) and only Haiku 4.5 still uses budget_tokens. Effort levels are low | medium | high (default) | xhigh – use xhigh for the hardest coding/agentic work on Opus 5 / Fable 5.1. On Fable 5.1 thinking is always on.

Thinking blocks and fallback

Thinking blocks are readable only by the producing model or a newer one. If you fall back from Fable 5.1 to an older model, the earlier thinking blocks are silently dropped – design for this in Domain 5’s fallback discussion.


3.6 Prefilling

Prefill the start of the assistant turn to constrain the shape of the output – force JSON, skip preamble, or lock a format.

python
messages = [
{"role": "user", "content": "Return the extracted fields."},
{"role": "assistant", "content": "{"}, # prefill: forces the model to continue JSON
]

Prefilling is a lightweight steer, not a guarantee. For strong guarantees use structured outputs or strict tools (next). Note: prefilling interacts poorly with extended thinking on models where thinking must come first – prefer structured outputs there.


3.7 Structured output: the three routes

This is the highest-value decision table in the domain.

RouteHowGuaranteeBest forCaveats
output_config.format (structured outputs)Pass a JSON schema in output_config.formatModel output conforms to the schemaYou want a typed response object, no tool semanticsNewer capability; check model support
Tool-use-as-schemaDefine a tool whose input_schema is your target shape; read the tool call’s inputStrong when combined with strict: trueYou already use tools, or want the model to “emit” a recordHistorically paired with forced tool_choice – restricted on Fable 5.1
Strict tools (strict: true)Add strict: true to a tool schemaEnforces the schema on the tool inputReliable tool arguments / record emissionPreferred on Fable 5.1 where forcing tool_choice is blocked
python
# Route 1: structured outputs via output_config.format
resp = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[{"role": "user", "content": f"<document>{doc}</document> Extract fields."}],
output_config={"format": {"type": "json_schema", "schema": INVOICE_SCHEMA}},
)
python
# Route 3: strict tool as a schema (works on Fable 5.1 with tool_choice: auto)
tools = [{
"name": "record_invoice",
"description": "Record the extracted invoice fields.",
"strict": True,
"input_schema": INVOICE_SCHEMA,
}]
resp = client.messages.create(
model="claude-fable-5-1",
max_tokens=1024,
tools=tools,
tool_choice={"type": "auto"}, # NOT {"type":"tool"} or "any" on Fable 5.1
messages=[{"role": "user", "content": prompt}],
)

Fable 5.1 · no forced tool choice

On Claude Fable 5.1, tool_choice: "any" and forced {"type": "tool", "name": …} return 400. To get structured output on Fable 5.1, use one of: tool_choice: "auto" plus an instruction to call the tool, strict: true tool schemas, or structured outputs (output_config.format). This is a favourite distractor: any option that forces a tool on Fable 5.1 is wrong.

Exam signal

“Guaranteed to match a schema, no tool semantics needed” → structured outputs. “Already using tools / want the model to emit a record” → tool-use-as-schema with strict: true. “On Fable 5.1” → never force tool_choice; use auto + instruction, strict, or structured outputs.


3.8 Schema design

A good schema is self-documenting and constraining.

json
{
"type": "object",
"properties": {
"invoice_number": { "type": "string", "description": "The vendor invoice ID as printed." },
"total": { "type": "number", "description": "Grand total in the invoice currency." },
"currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] },
"due_date": { "type": ["string", "null"], "description": "ISO 8601 date, or null if absent." },
"line_items": { "type": "array", "items": { "type": "object" } }
},
"required": ["invoice_number", "total", "currency"],
"additionalProperties": false
}
  • enum constrains a field to known values (prevents free-text drift).
  • required forces presence; omit optional fields from it.
  • Nullable via "type": ["string", "null"] – model the “absent” case explicitly rather than inviting a hallucinated value.
  • description on every field – the model reads these; they are the cheapest accuracy lever.
  • additionalProperties: false keeps output tight.

3.9 Validation-retry loop and defensive parsing

Even with schemas, validate and be ready to retry with specific error feedback.

python
import json, jsonschema
def extract_with_retry(client, prompt, schema, max_attempts=3):
messages = [{"role": "user", "content": prompt}]
for attempt in range(max_attempts):
resp = client.messages.create(
model="claude-opus-5", max_tokens=1024, messages=messages,
output_config={"format": {"type": "json_schema", "schema": schema}},
)
if resp.stop_reason == "refusal":
return {"status": "refused"} # do not retry blindly
if resp.stop_reason == "max_tokens":
raise OutputTruncated() # raise max_tokens and retry
text = "".join(b.text for b in resp.content if b.type == "text")
try:
data = json.loads(text) # defensive parse
jsonschema.validate(data, schema)
return {"status": "ok", "data": data}
except (json.JSONDecodeError, jsonschema.ValidationError) as e:
messages.append({"role": "assistant", "content": text})
messages.append({"role": "user",
"content": f"That output failed validation: {e}. "
f"Return only valid JSON matching the schema."})
return {"status": "failed_validation"}

Key points: feed the specific error back so the model self-corrects; parse defensively (never eval); treat refusal and max_tokens as distinct outcomes, not validation failures.


3.10 Extraction pipelines for heterogeneous documents

When you extract from many document types (invoices, contracts, receipts), a single aggregate accuracy number hides a type that is failing badly.

Anti-pattern 10 · Aggregate metrics masking per-type failure

Reporting “94% overall extraction accuracy” can hide that contracts are at 60% while invoices are at 99%. Always measure per document type (per-segment metrics). The exam-correct evaluation slices the metric by type and gates on the worst type.

text
┌─► classify type ─┬─► invoice schema ─► validate ─► metrics[invoice]
Document ─► route ┤ ├─► contract schema ─► validate ─► metrics[contract]
└─────────────────┴─► receipt schema ─► validate ─► metrics[receipt]
per-type accuracy, not one aggregate number

Design: route by type → apply the type-specific schema → validate → record per-type metrics. Report and alert per type.


3.11 Handling refusal and max_tokens

stop_reasonMeaningCorrect handling
refusalModel declined (safety)Stop; do not “retry harder”; escalate or reframe legitimately; log
max_tokensOutput truncatedNot a completion; raise max_tokens, or chunk the task, then retry
end_turnModel finishedProceed
tool_useWants a toolRun tool, continue loop

Treating max_tokens output as a complete, parseable result is a silent-failure trap.


3.12 Evaluator-optimizer prompts and independence

For a generate-critique-refine loop, give the evaluator the explicit rubric and run it as an independent context – a fresh session, ideally a different model. Reusing the generating conversation to grade itself keeps the reasoning bias that produced the flaw.

Anti-pattern 9 · Same-session self-review

“Ask the model in the same chat whether its answer is correct” retains reasoning-context bias. Use a separate session / different model as the judge, scoring against the written rubric. This applies to LLM-as-judge evals too.


3.13 Schema design patterns for hard extraction cases

Beyond enums and nullable types, several schema patterns decide whether extraction is reliable on messy documents.

PatternJSON SchemaSolves
Discriminated uniononeOf with a const discriminator fieldOne tool/endpoint handling several document types cleanly
Bounded arrays"maxItems": N on line_itemsRunaway output and truncation on huge tables
Formatted strings"format": "date", "pattern": "^[A-Z]{3}$"Dates/codes that must match a shape
Confidence & provenanceconfidence enum + source_span fieldDownstream gating and human verification
Explicit “not found”nullable + a found: boolean siblingDistinguishing “absent” from “missed”
json
{
"type": "object",
"oneOf": [
{ "properties": { "doc_type": { "const": "invoice" }, "total": { "type": "number" } },
"required": ["doc_type", "total"] },
{ "properties": { "doc_type": { "const": "contract" }, "term_months": { "type": "integer" } },
"required": ["doc_type", "term_months"] }
]
}

Exam signal

“The optional field is sometimes hallucinated when absent” → model it as nullable with an explicit found flag, not as a plain optional the model feels pressure to fill. “One pipeline, several document types” → discriminated union, not one loose object.


3.14 A robust validation-retry harness (with schema-error feedback)

The exam-correct loop treats refusal and max_tokens as distinct outcomes, feeds the specific validation error back, and never parses unsafely.

python
import json, jsonschema
from jsonschema import Draft202012Validator
def extract(client, prompt, schema, model="claude-opus-5", max_attempts=3):
messages = [{"role": "user", "content": prompt}]
for attempt in range(max_attempts):
resp = client.messages.create(
model=model, max_tokens=2048, messages=messages,
output_config={"format": {"type": "json_schema", "schema": schema}},
)
if resp.stop_reason == "refusal":
return {"status": "refused"} # safety stop — do not bypass
if resp.stop_reason == "max_tokens":
return {"status": "truncated", "retry": "raise_limit_or_chunk"} # not complete
text = "".join(b.text for b in resp.content if b.type == "text")
try:
data = json.loads(text) # never eval()
errors = sorted(Draft202012Validator(schema).iter_errors(data),
key=lambda e: e.path)
if not errors:
return {"status": "ok", "data": data}
detail = "; ".join(f"{list(e.path)}: {e.message}" for e in errors[:5])
except json.JSONDecodeError as e:
detail = f"invalid JSON at pos {e.pos}: {e.msg}"
messages += [
{"role": "assistant", "content": text},
{"role": "user", "content": f"Validation failed: {detail}. "
f"Return ONLY valid JSON matching the schema."},
]
return {"status": "failed_validation"}

Never eval model output

Parsing model output with eval() (Python) or Function/eval (JS) executes arbitrary code from an untrusted source. Always json.loads / JSON.parse behind a schema validator.


3.15 Structured-output cost and token arithmetic

Structured output is not free of token cost, and schema verbosity shows up on the bill. A quick model:

text
Prompt = system(1,500) + tools/schema(800) + document(6,000) = 8,300 input tokens
Output = ~400 tokens of JSON
On Opus 5 ($5/M in, $25/M out):
input = 8,300 × $5 / 1e6 = $0.0415
output = 400 × $25 / 1e6 = $0.0100
per doc ≈ $0.0515
Cache the stable 2,300-token prefix (system + schema), read ≈ 0.1×:
cached prefix = 2,300 × $0.5 / 1e6 = $0.00115 (vs $0.0115 uncached)
→ saves ≈ $0.0104 per doc; over 1M docs ≈ $10,400 saved.
Batch API (latency-tolerant) halves the remaining cost → ~$0.026/doc.
LeverEffectWhen correct
Cache the schema + system prefix~10× cheaper reads on the stable partRepeated extraction with a fixed schema
Batch API50% off, ≤24hLatency-tolerant bulk extraction
Cheaper model (Haiku 4.5)Lower per-token priceSimple, well-constrained schemas
Tighter schema (maxItems, closed objects)Fewer output tokens, less truncationLarge tables/line items

Exam signal

“MOST cost-effective way to extract from a million documents overnight” combines caching the stable schema/system prefix with the Batch API, and possibly a cheaper model — not “use a bigger model” or “call in real time at high concurrency”.


Common misconceptions

MisconceptionRealityWhy it matters on the exam
“Prefilling guarantees valid JSON.”Prefill only steers; structured outputs / strict tools guarantee the schema.Route-selection items.
“Force tool_choice for reliable tool output.”On Fable 5.1 forcing tool choice is a 400; use auto+instruction, strict, or structured outputs.The signature Fable 5.1 distractor.
“max_tokens output is just a bit shorter.”It is truncated; parsing it as complete corrupts data.Silent-truncation trap.
“A refusal is a validation error to retry.”Refusal is a safety stop; retrying to bypass is wrong.stop_reason handling items.
“One aggregate accuracy number is enough.”Aggregate hides a failing document type; slice per type (#10).Evaluation items.
“The model can grade its own answer in-session.”Same-session self-review inherits the bias (#9); use an independent judge.Evaluator-optimizer / LLM-as-judge items.
“budget_tokens controls thinking everywhere.”Only Haiku 4.5 uses budget_tokens; Opus/Sonnet/Fable use adaptive thinking + effort.Model-facts distractor.
“Variable input first is fine.”The stable prefix must be first to be cacheable.Caching items.

Scenario walkthrough — an extraction pipeline that looks fine but ships bad contracts

Situation. A pipeline extracts fields from invoices, contracts and receipts on Fable 5.1 and writes to a database. It reports 96% overall accuracy, yet the legal team keeps finding wrong contract term lengths. The engineers force tool_choice onto a record tool and see intermittent 400s; when a big contract is processed, the JSON sometimes ends mid-array and the code writes the partial object. A proposal is to “grade outputs by asking the model, in the same chat, if it is confident.”

Expert reasoning trace.

  1. Fix the 400s first. Fable 5.1 forbids forced tool_choice. Switch to structured outputs (output_config.format) or a strict: true tool with tool_choice: "auto" plus an instruction. Reject “retry with backoff” (a 400 is deterministic) and "any" (also blocked).

  2. Fix the silent truncation. JSON ending mid-array with stop_reason == max_tokens is truncation, not a result. Raise max_tokens, or chunk the contract (or bound line_items with maxItems), then retry. Never write the partial object.

  3. Fix the metric. 96% overall masks a failing type (#10). Report per-document-type accuracy, alert on contracts, and gate release on the worst type. Reject “bigger sample” and “average more runs” — both still aggregate.

  4. Fix the evaluation design. Same-session self-confidence is #9 and also self-reported confidence territory. Use an independent judge (fresh session, ideally a different model) scoring against a written rubric.

  5. Harden the schema. Contracts need a discriminated union branch with term_months as a typed integer and a nullable + found flag for optional clauses so absence is not hallucinated.

Exam-correct decision: structured outputs / strict tools (not forced choice), truncation handling by raising/chunking, per-type metrics gating on the worst type, an independent evaluator, and a discriminated-union schema. Each rejected option is a named trap (Fable forced choice, silent truncation, aggregate masking #10, same-session #9).


Exam traps in this domain

TrapWhy it is wrong
Force tool_choice: "any"/{"type":"tool"} on Fable 5.1Returns 400; use auto+instruction, strict, or structured outputs
Put the variable task before the stable system promptBreaks caching; stable content must be first
Use one aggregate accuracy metric across document typesHides per-type failure (anti-pattern 10)
Treat max_tokens output as a complete resultIt is truncated; raise the limit or chunk
Retry a refusal by rephrasing to bypass safetyRefusals are safety stops, not validation errors
Grade an answer in the same session that produced itSame-session self-review bias (anti-pattern 9)
Set budget_tokens on Opus 5 / Sonnet 5 / Fable 5.1Removed (400); only Haiku 4.5 uses it
Rely on prefilling as a hard guaranteeIt steers, not guarantees; use structured outputs/strict tools
Skip field descriptions in the schemaDescriptions are the cheapest accuracy lever
Edit earlier turns mid-session on Fable 5.1Invalidates thinking blocks; harness must be append-only
Parse model output with eval()Executes untrusted code; use json.loads/JSON.parse behind a validator
Model an optional field as a plain optionalThe model may hallucinate a value; use nullable + a found flag
Use one loose object for several document typesUse a discriminated union (oneOf + const)
Leave line-item arrays unboundedAdd maxItems to prevent runaway output and truncation
Real-time high concurrency for bulk overnight extractionUse the Batch API (50% off, ≤24h) and cache the schema prefix

Practice questions

Q1 · A pipeline on Claude Fable 5.1 must return records that match a JSON schema. An engineer sets a forced tool_choice on the record tool and gets 400 errors. What is the correct fix? (Select one)

A. Retry with exponential backoff. B. Use tool_choice: 'auto' with an instruction to call the tool, or strict: true tools, or structured outputs via output_config.format. C. Switch to tool_choice: 'any'. D. Lower max_tokens.

Answer: B. Fable 5.1 rejects forced tool choice (any and {type:"tool"} both 400). The supported routes are auto+instruction, strict tools, or structured outputs. Backoff (A) does not fix a 400; any (C) is also blocked; max_tokens (D) is unrelated.

Q2 · A team reports 94% overall extraction accuracy across invoices, contracts and receipts. A stakeholder complains contracts are frequently wrong. What is the BEST evaluation change? (Select one)

A. Increase the overall sample size. B. Report per-document-type accuracy and gate on the worst-performing type. C. Raise the model temperature for contracts. D. Average three runs to smooth the metric.

Answer: B. Aggregate accuracy masks a failing type (anti-pattern 10). Per-type metrics expose and gate on the weak type. Larger samples (A) still aggregate; temperature (C) does not fix accuracy; averaging (D) further hides the problem.

Q3 · A prompt places the specific user document first and the long stable system rules last. Latency and cost are high due to no cache hits. What is the fix? (Select one)

A. Shorten the document. B. Put the stable system prompt, tools and reference material first with cache_control on the last stable block, and the variable task after. C. Disable caching. D. Use a bigger model.

Answer: B. Caching requires the stable content first; the variable task goes last. Shortening the document (A) does not enable caching, disabling caching (C) worsens cost, and model size (D) is irrelevant.

Q4 · An extraction returns text ending mid-object and `stop_reason` is `max_tokens`. The code JSON-parses it and fails. What is the correct handling? (Select one)

A. Treat the partial JSON as the result. B. Recognise max_tokens as truncation, raise the token limit or chunk the task, then retry — do not parse it as complete. C. Log a generic error and return empty. D. Ask the model in the same chat if it is sure.

Answer: B. max_tokens means truncated output, not completion. Raise the limit or chunk. Parsing partial output (A) is a silent-failure trap; returning empty (C) is anti-pattern 7; same-session self-check (D) is unrelated.

Q5 · Which TWO schema design choices most improve extraction reliability? (Select two)

A. A description on every field. B. Free-text for fields that have a fixed set of values. C. An enum for the currency field and explicit nullable types for optional fields. D. Omitting required entirely. E. Allowing additionalProperties: true.

Answer: A and C. Descriptions guide the model and enums/nullable types constrain and model absence explicitly. Free text (B) invites drift, omitting required (D) loses presence guarantees, and open additional properties (E) loosens the output.

Q6 · A validation-retry loop keeps failing. Currently it just re-sends the same prompt. What change most helps the model self-correct? (Select one)

A. Increase the number of retries only. B. Feed the specific validation error back to the model and ask it to return only valid JSON matching the schema. C. Switch to eval() to parse the output. D. Lower max_tokens.

Answer: B. Specific error feedback lets the model fix the exact problem. More blind retries (A) rarely help, eval() (C) is unsafe, and lowering max_tokens (D) risks truncation.

Q7 · When should you use extended thinking rather than prompted chain-of-thought? (Select one)

A. For every request, always. B. For hard multi-step reasoning or agentic planning, using thinking: {'type': 'adaptive'} on current models. C. Never; chain-of-thought is always better. D. Only to reduce cost.

Answer: B. Extended thinking suits hard, multi-step or agentic reasoning. It is not needed for every request (A), is not universally worse (C), and does not reduce cost (D) — thinking tokens add cost.

Q8 · An evaluator-optimizer loop grades the translation in the same conversation that produced it and quality does not improve. What is the flaw? (Select one)

A. The rubric is too detailed. B. Same-session self-review retains reasoning bias; run the evaluator as an independent context, ideally a different model, against the explicit rubric. C. The generator needs a higher temperature. D. There are too few retries.

Answer: B. Anti-pattern 9. Independence removes the shared bias. A detailed rubric (A) is good, temperature (C) does not fix bias, and more retries (D) repeat the biased judge.

Q9 · A document contains untrusted third-party text that instructs the model to ignore its task. What prompt-design practice reduces this risk? (Select one)

A. Concatenate the document directly into the instructions. B. Wrap the untrusted content in a clearly named XML tag (e.g. <document>…</document>) and instruct the model to treat it as data, not instructions. C. Trust the model to notice. D. Set temperature to 0.

Answer: B. XML content boundaries separate data from instructions and blunt indirect injection. Concatenation (A) invites injection, trusting the model (C) is not a control, and temperature (D) is irrelevant.

Q10 · On Claude Sonnet 5, a harness tries to inject a mid-conversation system message. What is true? (Select one)

A. Sonnet 5 allows mid-conversation system messages. B. Sonnet 5 does not allow mid-conversation system messages; design the system prompt up front. C. Only Fable 5.1 forbids this. D. Set budget_tokens to enable it.

Answer: B. Sonnet 5 does not allow mid-conversation system messages, so the design must fix the system prompt up front. A is false; Fable 5.1’s append-only constraint is separate (C); budget_tokens (D) is removed on Sonnet 5.

Q11 · A pipeline needs a guaranteed schema-conformant object and does not use tools for anything else. Which route is cleanest? (Select one)

A. Prefill the assistant turn with { and hope. B. Structured outputs via output_config.format with a JSON schema. C. Force tool_choice to a dummy tool on Fable 5.1. D. Parse free-text prose with a regex.

Answer: B. Structured outputs give a schema guarantee without tool semantics. Prefill (A) only steers, forcing tool_choice on Fable 5.1 (C) 400s, and regex on prose (D) is brittle.

Q13 · One extraction pipeline must handle invoices, contracts and receipts, each with different required fields. Which schema pattern is BEST? (Select one)

A. One loose object with every possible field optional. B. A discriminated union (oneOf with a const doc_type discriminator), so each branch enforces its own required fields. C. Three unrelated endpoints with no shared contract. D. A single string field holding raw JSON text.

Answer: B. A discriminated union cleanly enforces per-type requirements in one schema. A loosens everything and invites drift; C loses a shared contract; D abandons schema guarantees entirely.

Q14 · An optional 'renewal_clause' field is frequently hallucinated when the clause is absent. Which schema change MOST reduces this? (Select one)

A. Make it a required string so the model always fills it. B. Model it as nullable ([\"string\", \"null\"]) with an explicit renewal_found: boolean sibling, so absence is represented rather than invented. C. Remove the field entirely. D. Raise the temperature so answers vary.

Answer: B. Nullable plus an explicit found flag lets the model represent absence instead of hallucinating. Required (A) forces a value; removing it (C) loses the data; temperature (D) worsens consistency.

Q15 · A million documents must be extracted overnight as cheaply as possible with a fixed schema. Which combination is MOST cost-effective? (Select two)

A. Cache the stable system + schema prefix so reads cost ~0.1x. B. Use the Message Batches API for the latency-tolerant bulk run (50% off, ≤24h). C. Call the real-time API at maximum concurrency. D. Use Fable 5.1 at $10/$50 for every document. E. Randomise the prompt each call to avoid stale reads.

Answer: A and B. Caching the fixed prefix and the Batch API together cut cost sharply for latency-tolerant bulk work. Real-time high concurrency (C) costs full price and risks limits; the most expensive model (D) raises cost; randomising (E) destroys the cacheable prefix.

Q16 · A developer parses extraction output with `eval()` because 'it handles trailing commas'. What is the correct critique? (Select one)

A. It is fine if the model is trusted. B. eval() executes arbitrary untrusted code from the model; parse with json.loads/JSON.parse behind a schema validator and feed validation errors back on failure. C. Use a regex instead of eval(). D. Lower max_tokens so the output is smaller.

Answer: B. Model output is untrusted; eval() is a code-execution risk. Safe parsing plus schema validation and error feedback is correct. The model is never ‘trusted’ for eval (A); a regex (C) is brittle; max_tokens (D) is unrelated.

Q17 · A large contract's extracted JSON ends mid-array with `stop_reason` `max_tokens`, and the pipeline writes the partial object. What is the correct handling? (Select one)

A. Write the partial object; it is mostly complete. B. Treat max_tokens as truncation: raise the limit or chunk the document (or bound arrays with maxItems), then retry — never persist the partial output. C. Return an empty object so the pipeline continues. D. Ask the model in the same session whether it finished.

Answer: B. max_tokens is truncation, not completion. Persisting the partial (A) corrupts data; returning empty (C) is silent failure (#7); same-session self-check (D) does not address truncation.

Q18 · On Opus 5 an agentic harness sets `budget_tokens` for thinking and receives a 400. What is true and what is the fix? (Select one)

A. budget_tokens works on Opus 5; retry the 400. B. budget_tokens is Haiku-4.5-only; on Opus 5 use adaptive thinking with an effort level (low|medium|high|xhigh), e.g. xhigh for the hardest work. C. Disable thinking to avoid the error. D. Switch to tool_choice: any to enable budgets.

Answer: B. Only Haiku 4.5 uses budget_tokens; Opus 5 uses effort levels. The 400 is deterministic, not transient (A); disabling thinking (C) removes needed reasoning; tool_choice (D) is unrelated to thinking control.

Key takeaways

  • Design prompts as versioned, modular systems: stable content first (system role, tools, docs) with cache_control; variable task last.
  • State explicit success criteria and rubrics; use XML tags to separate untrusted data from instructions.
  • Pick structured output deliberately: output_config.format for schema guarantees, tool-use-as-schema/strict for record emission; never force tool_choice on Fable 5.1.
  • Design schemas with enums, required, nullable types, descriptions, discriminated unions and bounded arrays; model absence with nullable + a found flag.
  • Build a validation-retry loop that feeds the specific schema error back and parses safely (never eval); treat refusal and max_tokens as distinct outcomes.
  • Slice extraction metrics per document type and gate on the worst; never trust one aggregate number.
  • Control reasoning with adaptive thinking + effort (Opus/Sonnet/Fable); only Haiku 4.5 uses budget_tokens.
  • Run evaluators as independent contexts; on Fable 5.1 keep the harness append-only and freeze system/tools.
  • For bulk extraction, cache the stable schema/system prefix and use the Batch API (50% off) — the cost levers, not a bigger model.

Last updated Sep 18, 2026