# D8 · Eval, Testing and Debugging

Identifying integration-layer vs model-output errors, recovery, trace analysis, request-ID logging, reproducibility, eval basics, regression tests in CI and per-segment metrics.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This domain is roughly **1 of 53 items**, but it ties the others together. It tests whether you can tell an **integration-layer** failure from a **model-output** failure, reproduce and trace problems, and evaluate quality without fooling yourself. The theme: **measure per-segment, reproduce deterministically, and judge output in a separate session.**

## Learning objectives

By the end of this page you should be able to:

1. Distinguish **integration-layer** from **model-output** errors and choose recovery.
2. Do **trace analysis** and log **request IDs**.
3. **Reproduce** issues with a pinned model and `temperature: 0` where appropriate.
4. Apply **eval basics**: golden set, rubric, LLM-as-judge in a separate session.
5. Add **regression tests in CI** and use **per-segment metrics**.

---

## 8.1 Integration-layer vs model-output errors

The first question in any incident: **which layer failed?** Integration-layer errors live in the plumbing (auth, transport, request shape, rate limits, parsing). Model-output errors live in what Claude produced (hallucination, format drift, refusal, truncation). Retries and timeouts fix the former; prompt/context/eval changes fix the latter. Applying the wrong fix to the wrong layer is the classic distractor.

### Error taxonomy

| Layer | Error | Diagnostic signal | Recovery |
| --- | --- | --- | --- |
| Integration | **Auth** | `401 authentication` / `403 permission` | Check API key, org, scopes; do not retry blindly |
| Integration | **Schema / bad request** | `400 invalid_request` (e.g. `tool_choice:"any"` on Fable 5.1) | Fix request shape; use `auto` + `strict:true` / structured outputs |
| Integration | **Timeout** | Client timeout, no response | Set sane timeouts; retry with backoff; consider streaming |
| Integration | **Rate limit** | `429 rate_limit`, `retry-after` header | Exponential backoff + jitter; respect `retry-after`; batch/queue |
| Integration | **Parsing** | JSON decode error on the response body | Validate/repair; request structured outputs; validation-retry |
| Model output | **Hallucination** | Fluent claim with no grounding in inputs | Grounding, citations, RAG, validation; verify claims |
| Model output | **Format drift** | Output nearly-valid but off-schema | `strict:true` tool schema / structured outputs; validation-retry |
| Model output | **Refusal** | `stop_reason: "refusal"` | Reframe legitimate request; adjust system prompt; escalate |
| Model output | **Truncation** | `stop_reason: "max_tokens"` | Raise `max_tokens`; chunk the task; stream and continue |

:::tip[Exam signal]
First ask **which layer**. A `429`, a timeout, or a JSON decode error is integration — fix with backoff, timeouts, or request/schema changes. A hallucination, off-schema output, or refusal is model output — fix with prompt/context/model changes and validation. `stop_reason` and the HTTP status code tell you which layer.
:::

---

## 8.2 Trace analysis walkthrough

Log the **`request-id`** and a correlation ID for every call, plus `model`, `stop_reason`, `usage` and latency. For agents, log the **tool sequence**. Traces let you pinpoint where a multi-step run went wrong and give Anthropic support a handle.

```python
resp = client.messages.create(model=MODEL, max_tokens=512, messages=msgs)
log.info("claude_call", extra={
    "request_id": resp._request_id, "model": resp.model,
    "stop_reason": resp.stop_reason, "in": resp.usage.input_tokens,
    "out": resp.usage.output_tokens})
```

Reading a trace of a stuck agent:

```text
corr_id=abc-123
 step 1  request_id=req_01A  stop_reason=tool_use   tools=[search_orders]        in=1,240 out=180
 step 2  request_id=req_01B  stop_reason=tool_use   tools=[get_order(id=9981)]   in=1,910 out=95
 step 3  request_id=req_01C  stop_reason=tool_use   tools=[get_order(id=9981)]   in=2,600 out=95   ← repeat
 step 4  request_id=req_01D  stop_reason=tool_use   tools=[get_order(id=9981)]   in=3,290 out=95   ← loop
 ...
 step N  request_id=req_01Z  stop_reason=max_tokens                              in=9,900 out=512  ← truncated
```

Diagnosis: the loop repeats the same tool call and input tokens climb each turn — the harness is not feeding the tool **result** back, or termination is driven by something other than `stop_reason`. This is anti-patterns #1/#2 (natural-language termination and iteration caps) rather than a model defect. The `max_tokens` at the end is a downstream symptom, not the cause. Fix the loop: after each `tool_use`, append the tool result and continue until `stop_reason` is `end_turn`.

---

## 8.3 Reproducibility

To reproduce a model-output issue, control every variable you can:

1. **Pin the exact model snapshot** — not a floating alias.
2. **Set `temperature: 0`** where appropriate to minimise variance.
3. **Freeze the system prompt and tools** and replay the **same messages** in order.

```python
resp = client.messages.create(
    model="claude-sonnet-5",   # pinned snapshot
    temperature=0,             # minimise variance
    max_tokens=512,
    system=FROZEN_SYSTEM_PROMPT,
    tools=FROZEN_TOOLS,
    messages=RECORDED_MESSAGES,  # exact replay
)
```

Non-determinism means the output still is not byte-identical, but pinning + `temperature: 0` + a frozen prompt makes issues far more repeatable and comparisons fair. Note: Fable 5.x / Opus 5 / Sonnet 5 harnesses must be **append-only** — editing or reordering earlier turns invalidates later thinking blocks, so replay by appending, never by rewriting history.

---

## 8.4 Eval basics and an eval harness

| Element | What it is |
| --- | --- |
| **Golden set** | Curated inputs with known-good outputs |
| **Exact match** | For deterministic outputs (labels, extractions) |
| **Rubric** | Explicit scoring criteria for open-ended output |
| **LLM-as-judge** | A *different* model/session scores output against the rubric |

:::danger[Same-session self-review]
Never have the model grade its own output in the same session — it retains the reasoning bias that produced it (anti-pattern #9). Use a separate session and ideally a different model as judge.
:::

A minimal harness combining exact-match and an LLM-as-judge in a **separate call**:

```python
import json
from anthropic import Anthropic

client = Anthropic()
GEN_MODEL = "claude-sonnet-5"      # system under test
JUDGE_MODEL = "claude-opus-5"      # different model, separate call

def generate(case):
    r = client.messages.create(
        model=GEN_MODEL, temperature=0, max_tokens=512,
        system="Extract the invoice total as JSON: {'total_cents': int}.",
        messages=[{"role": "user", "content": case["input"]}],
    )
    return r.content[0].text

def exact_match(pred, expected):
    try:
        return json.loads(pred).get("total_cents") == expected["total_cents"]
    except json.JSONDecodeError:
        return False  # parsing failure counts as a miss, never silently passed

RUBRIC = (
    "Score 1 if the answer is faithful to the source and correctly formatted, "
    "else 0. Return JSON: {'score': 0 or 1, 'reason': str}."
)

def llm_judge(case, pred):
    # Separate session/model — no shared context with the generator.
    r = client.messages.create(
        model=JUDGE_MODEL, temperature=0, max_tokens=256,
        system=RUBRIC,
        messages=[{"role": "user",
                   "content": f"SOURCE:\n{case['input']}\n\nANSWER:\n{pred}"}],
    )
    return json.loads(r.content[0].text)

def run(golden):
    results = []
    for case in golden:
        pred = generate(case)
        results.append({
            "id": case["id"], "segment": case["segment"],
            "exact": exact_match(pred, case["expected"]),
            "judge": llm_judge(case, pred)["score"],
        })
    return results
```

---

## 8.5 Regression tests in CI

Run the golden set in **CI** on every prompt/model/config change and fail the build on regression. Pin the model and `temperature: 0` for stable comparisons.

```yaml
name: eval-regression
on:
  pull_request:
    paths: ["prompts/**", "src/**", "evals/**"]
jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install anthropic
      - name: Run golden-set eval
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
        run: python evals/run.py --golden evals/golden.jsonl --min-score 0.95
      # run.py exits non-zero if aggregate OR any per-segment score drops below threshold
```

---

## 8.6 Per-segment metrics

Report metrics **per segment** (per document type, per language, per intent). **Aggregate accuracy masks per-segment failures** (anti-pattern #10) — the single most common eval trap on the exam.

| Segment | Cases | Accuracy | Verdict |
| --- | --- | --- | --- |
| **Overall (aggregate)** | 1,000 | **92%** | Looks fine — misleading |
| invoices | 600 | 98% | Healthy |
| receipts | 300 | 95% | Healthy |
| handwritten | 100 | **61%** | The real problem, hidden by the aggregate |

The 92% aggregate passes a naive gate, yet handwritten forms fail 4 in 10. Gate CI on the **worst segment**, not the mean, and always break metrics down before shipping.

---

## 8.7 Choosing the right eval method for the output type

Not every output is scored the same way. Picking the wrong scoring method is a subtle exam trap.

| Output type | Best scoring method | Why |
| --- | --- | --- |
| Labels / classifications | **Exact match** (accuracy, F1 per class) | Deterministic ground truth |
| Structured extraction | **Field-level match** against expected JSON | Partial credit and per-field diagnosis |
| Open-ended prose | **Rubric + LLM-as-judge** (separate session) | No single correct string |
| Comparative quality | **Pairwise A/B** (which of two is better) | Humans/judges compare more reliably than absolute-score |
| Retrieval/grounding | **Citation/faithfulness checks** | Catches ungrounded claims |

```python
# Pairwise A/B: ask a separate judge which answer is better, randomising order to avoid position bias.
import random

def pairwise(judge_client, question, ans_a, ans_b):
    first, second = ("A", ans_a, "B", ans_b) if random.random() < 0.5 else ("B", ans_b, "A", ans_a)
    label1, text1, label2, text2 = first
    r = judge_client.messages.create(
        model="claude-opus-5", max_tokens=128, temperature=0,
        system="Reply with the label of the better answer: 'A' or 'B'. JSON: {'winner': 'A'|'B'}.",
        messages=[{"role": "user",
                   "content": f"Q:{question}\n\n{label1}:\n{text1}\n\n{label2}:\n{text2}"}])
    import json
    return json.loads("".join(b.text for b in r.content if b.type == "text"))["winner"]
```

:::tip[Exam signal]
'Deterministic labels' → exact match. 'Open-ended quality' → rubric + LLM-as-judge in a separate session. 'Which version is better' → pairwise A/B. 'Grounding' → citation/faithfulness. Match the method to the output type.
:::

---

## 8.8 Debugging decision tree: which layer, which fix

A fast triage flow turns a vague 'it's broken' into a targeted fix. The HTTP status and `stop_reason` are your first two signals.

```text
Failure observed
   │
   ├─ HTTP 4xx/5xx?  ──► INTEGRATION layer
   │     ├─ 429 / 500 / 529 ──► transient: backoff + jitter, honour retry-after
   │     ├─ 400 invalid_request ──► fix request shape (e.g. forced tool on Fable 5.1)
   │     ├─ 401 / 403 ──► auth/permissions: fix key/org/scopes (do NOT retry)
   │     └─ timeout / JSON decode ──► timeouts/streaming; schema + validation-retry
   │
   └─ HTTP 200 but bad answer?  ──► MODEL-OUTPUT layer
         ├─ stop_reason == max_tokens ──► truncation: continue / raise cap / chunk
         ├─ stop_reason == refusal ──► safety decline: reframe / policy path
         ├─ off-schema-but-close ──► strict/structured outputs + validation-retry
         └─ confident but wrong ──► grounding: docs-first, citations, RAG, claim validation
```

| Symptom | Layer | Wrong fix (distractor) | Right fix |
| --- | --- | --- | --- |
| `429` under load | Integration | Rewrite the prompt | Backoff + `retry-after`; batch/queue |
| `400` on forced tool (Fable 5.1) | Integration | Retry with backoff | `auto`/structured outputs |
| Truncated JSON, `max_tokens` | Model output | Rotate API key | Raise cap / continue / chunk |
| Fabricated value | Model output | Raise `temperature` | Grounding + claim validation |

:::tip[Exam signal]
Read the **status code** and **`stop_reason`** first. A `429`/timeout/JSON error is integration; a hallucination/off-schema/refusal is model output. Applying an integration fix to a model-output problem (or vice-versa) is the classic distractor.
:::

---

## 8.9 Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| Any failure can be retried | Only transient integration errors (`429`/`5xx`/`529`) are retryable | Wrong-layer fix |
| A `400` might be transient | It is deterministic; fix the request | Distinguishes layers |
| Aggregate accuracy is enough | It masks per-segment failures (anti-pattern #10) | The top eval trap |
| Same-session self-review is efficient | The judge inherits the generator's bias (anti-pattern #9) | Judge-design trap |
| `temperature: 0` gives reproducible bytes | It reduces variance; pin the snapshot too | Reproducibility nuance |
| A `max_tokens` end is the root cause of a loop | It is a downstream symptom; fix loop/`stop_reason` handling | Trace-reading trap |
| Empty output means 'no results' | It may be silently suppressed error (anti-pattern #7) | Silent-failure trap |
| Exact match fits every output | Open-ended output needs a rubric/LLM-judge; comparisons need pairwise | Eval-method selection |

---

## 8.10 Scenario walkthrough: an eval that lied

**Scenario.** A document-classification service reports **93% accuracy** on its golden set and passes CI, yet production incidents keep involving **handwritten forms** and **non-English documents**. Investigation reveals: (1) CI gates only on aggregate accuracy; (2) the eval had the *generator* model grade its own open-ended rationales in the same session; (3) an intermittent `400 invalid_request` in production is being retried with exponential backoff and never succeeding; and (4) a helper returns an empty list when a call throws, which the caller treats as 'no matches'. Fix the eval and the debugging.

**Expert reasoning trace.**

1. **Break the aggregate.** 93% overall hides that handwritten and non-English segments fail badly (**anti-pattern #10**). Report **per-segment metrics** (by document type and language) and **gate CI on the worst segment**, not the mean.
2. **Fix the judge.** Same-session self-grading (**anti-pattern #9**) rubber-stamps the generator. Move to an **LLM-as-judge in a separate session, ideally a different model**, against an explicit rubric; for the open-ended rationales, consider **pairwise A/B** with randomised order.
3. **Fix the `400` triage.** A `400 invalid_request` is an **integration/schema** error and **deterministic** — retrying with backoff loops forever. Read the error: likely a forbidden parameter (e.g. forced `tool_choice` on Fable 5.1). Fix the request shape (`auto`/structured outputs); do not retry.
4. **Fix the silent suppression.** Returning an empty list on exception (**anti-pattern #7**) reports failure as success. Propagate the error with diagnostic context, or fail loudly; never let empty-as-success flow downstream.
5. **Make it reproducible.** Pin the model snapshot and set `temperature: 0` for the eval; log `request_id`, `stop_reason`, `usage` and latency so incidents are traceable.
6. **Reject the tempting alternatives.** 'Raise the aggregate threshold to 95%' — still masks per-segment failures. 'Retry the 400 more aggressively' — deterministic error. 'Trust the empty result as no matches' — silent suppression. 'Have the same model re-grade to save cost' — reintroduces bias.

**Correct decision.** Per-segment metrics with worst-segment CI gating; a separate-session/different-model LLM-judge (pairwise A/B for open-ended); fix the `400` at the request layer (no backoff); surface the suppressed error loudly; pin snapshot + `temperature: 0` and log request IDs for traceability.

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| Retrying a model-output error like a network error | Wrong layer; fix the prompt/eval, not backoff |
| Treating a `400` schema error as transient | It is deterministic; fix the request shape, do not retry |
| Same-session self-review | Anti-pattern #9; use a separate judge session/model |
| Reporting only aggregate accuracy | Anti-pattern #10; masks per-segment failures |
| Not logging request IDs | No trace handle for debugging/support |
| Reproducing with random temperature and unpinned model | Not reproducible; pin snapshot + temperature 0 |
| Treating empty output as success | Silent suppression (anti-pattern #7) |
| No CI regression suite | Prompt/model changes silently regress |
| Judging with the same model in the same context | Bias; use a different session/model |
| Reading `max_tokens` truncation as the root cause of a loop | It is a downstream symptom; fix the loop/`stop_reason` handling |
| Raising the aggregate threshold instead of gating per-segment | Still masks the failing segment; gate on the worst segment |
| Using exact match for open-ended output | Use a rubric + LLM-as-judge (separate session); pairwise A/B for comparisons |
| Not reading the status code and `stop_reason` before choosing a fix | They tell you the layer; the wrong-layer fix is the classic distractor |
| Ignoring position bias in pairwise judging | Randomise A/B order so the judge is not swayed by position |
| Treating a deterministic `400` as flaky and retrying | Fix the request shape; backoff will loop forever |

---

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · A service intermittently fails with JSON decode errors and 429s, and separately sometimes extracts the wrong invoice total. How should the team triage? (Select one)">
    A. Treat both as model problems and rewrite the prompt.
    B. Separate layers: fix 429s with backoff and JSON errors with validation-retry (integration); fix wrong extraction with prompt/context engineering and validation (model output).
    C. Treat both as network problems and add retries.
    D. Increase `max_tokens` for both.

    **Answer: B.** The failures are in different layers and need different fixes. Blanket prompt rewrites (A) ignore the integration errors; retries (C) do nothing for wrong extraction; raising `max_tokens` (D) addresses neither a 429 nor a wrong value.
  </AccordionItem>

  <AccordionItem title="Q2 · A model scores 91% overall on the eval, but a production incident involves handwritten forms. What eval practice would have surfaced this, and how should output be judged? (Select two)">
    A. Report per-segment metrics (by document type) instead of only aggregate.
    B. Have the same session grade its own output.
    C. Use an LLM-as-judge in a separate session/model against a rubric.
    D. Only track overall accuracy.
    E. Skip evals in CI.

    **Answer: A and C.** Per-segment metrics expose the handwritten-forms failure (anti-pattern #10), and a separate-session judge avoids self-review bias (anti-pattern #9). Aggregate-only (D), same-session grading (B) and no CI (E) are the anti-patterns.
  </AccordionItem>

  <AccordionItem title="Q3 · Every request to Claude Fable 5.1 returns `400 invalid_request`; the code sets `tool_choice: {'type':'any'}`. Which layer is this and what is the fix? (Select one)">
    A. Model-output error; rewrite the prompt to be clearer.
    B. Integration/schema error; Fable 5.1 rejects forced `any`/`{type:tool}`, so use `auto` with an instruction, `strict:true`, or structured outputs.
    C. Rate-limit error; add exponential backoff.
    D. Transient error; retry with jitter.

    **Answer: B.** A `400` on a forbidden parameter is a deterministic integration/schema error specific to Fable 5.1's tool-choice restriction; the fix is to change the request. It is not about prompt quality (A); it is not a `429` (C); and retrying a deterministic `400` (D) will always fail again.
  </AccordionItem>

  <AccordionItem title="Q4 · A trace shows an agent calling the same tool with identical input on every step, with input tokens climbing each turn, ending in `stop_reason: max_tokens`. What is the root cause? (Select one)">
    A. The model ran out of output tokens; raise `max_tokens`.
    B. The harness is not feeding the tool result back and/or termination is not driven by `stop_reason`; fix the loop to append results and stop on `end_turn`.
    C. Rate limiting; add backoff.
    D. A hallucination; add citations.

    **Answer: B.** Repeating the same call with growing context is a loop-control defect (anti-patterns #1/#2). The final `max_tokens` is a downstream symptom, not the cause (A). It is not a transport (C) or grounding (D) issue.
  </AccordionItem>

  <AccordionItem title="Q5 · A response comes back with `stop_reason: 'max_tokens'` and the JSON is cut off mid-object. Which layer, and what are TWO valid recoveries? (Select two)">
    A. Model-output truncation; raise `max_tokens` for the call.
    B. Chunk the task or stream and continue generation.
    C. It is an auth error; rotate the API key.
    D. Silently return the partial JSON as success.
    E. Retry unchanged with backoff.

    **Answer: A and B.** `max_tokens` means the output was truncated; raising the limit or splitting/streaming the work recovers it. It is not auth (C); returning partial output as success is silent suppression, anti-pattern #7 (D); retrying unchanged (E) truncates again.
  </AccordionItem>

  <AccordionItem title="Q6 · To reproduce a model-output bug reliably, which combination should the team use? (Select one)">
    A. Latest floating model alias, temperature 1, paraphrased prompt.
    B. Pinned model snapshot, `temperature: 0`, frozen system prompt and tools, exact message replay.
    C. Any model, as long as the prompt is similar.
    D. A different model each run to average out noise.

    **Answer: B.** Controlling the model snapshot, temperature, prompt, tools, and message history maximises repeatability. Floating aliases and non-zero temperature (A), loose prompts (C), and varying the model (D) all inject variance that defeats reproduction.
  </AccordionItem>

  <AccordionItem title="Q7 · An eval pipeline has the generator model grade its own answers in the same conversation. What is wrong, and what is the fix? (Select one)">
    A. Nothing; self-grading is efficient.
    B. Same-session self-review retains the reasoning bias that produced the answer; run an LLM-as-judge in a separate session, ideally a different model.
    C. Use exact match for everything instead.
    D. Grade only the aggregate.

    **Answer: B.** Same-session self-review (anti-pattern #9) is biased because the judge shares the generator's context. A separate-session, different-model judge removes that bias. Exact match (C) does not fit open-ended output; aggregate-only grading (D) is a different anti-pattern.
  </AccordionItem>

  <AccordionItem title="Q8 · Which artefacts should be logged for every Claude call to enable trace analysis? (Select two)">
    A. `request_id` and a correlation ID.
    B. The user's password.
    C. `model`, `stop_reason`, `usage` tokens, latency, and (for agents) the tool sequence.
    D. Nothing, to save storage.
    E. Only the final answer text.

    **Answer: A and C.** Request/correlation IDs plus model, stop reason, token usage, latency, and tool sequence give a full trace handle for debugging and support. Logging secrets (B) is a security violation; logging nothing (D) or only the answer (E) leaves you blind during incidents.
  </AccordionItem>

  <AccordionItem title="Q9 · A CI job runs the golden set but only fails the build if aggregate accuracy drops. A per-segment regression on 'legal documents' ships to production. What should CI do instead? (Select one)">
    A. Keep gating on aggregate; it is simpler.
    B. Gate on the worst-performing segment as well as the aggregate, failing if any segment drops below its threshold.
    C. Remove the eval from CI.
    D. Only run evals monthly.

    **Answer: B.** Gating on per-segment thresholds catches the failures an aggregate hides (anti-pattern #10). Aggregate-only gating (A) is exactly what let the regression through; removing (C) or slowing (D) evals makes it worse.
  </AccordionItem>

  <AccordionItem title="Q10 · A function returns an empty list when the Claude call raises an exception, and the caller treats empty as 'no results found'. Which anti-pattern is this and how is it fixed? (Select one)">
    A. Aggregate metrics; add per-segment reporting.
    B. Silent error suppression (anti-pattern #7); surface the error with diagnostic context instead of returning empty-as-success.
    C. Same-session self-review; use a separate judge.
    D. Iteration cap; drive from `stop_reason`.

    **Answer: B.** Returning empty on failure hides errors and reports failure as success — silent suppression (anti-pattern #7). The fix is to propagate the error with context. The other options name unrelated anti-patterns.
  </AccordionItem>

  <AccordionItem title="Q11 · Output is almost valid JSON but occasionally emits an extra trailing field not in the schema. Which layer, and what is the most robust fix? (Select one)">
    A. Integration timeout; add retries.
    B. Model-output format drift; enforce the schema with `strict: true` on the tool or structured outputs, plus a validation-retry.
    C. Rate limit; add backoff.
    D. Auth error; rotate the key.

    **Answer: B.** Off-schema-but-close output is format drift, a model-output problem; `strict:true`/structured outputs plus validation-retry enforce the schema. It is not a timeout (A), rate limit (C), or auth (D) issue.
  </AccordionItem>

  <AccordionItem title="Q12 · A team wants to prevent prompt and model changes from silently degrading quality. Which practice is essential? (Select one)">
    A. Manual spot-checks whenever someone remembers.
    B. A regression suite over a golden set run automatically in CI on every prompt/model/config change, with per-segment thresholds that fail the build.
    C. Trusting the model's self-reported confidence.
    D. Only measuring latency.

    **Answer: B.** Automated CI regression over a golden set with per-segment gating is the discipline that catches silent quality drops. Ad-hoc manual checks (A) are unreliable; self-reported confidence (C) is an anti-pattern; latency (D) does not measure quality.
  </AccordionItem>

  <AccordionItem title="Q13 · An eval scores open-ended rationales. Which scoring method is appropriate, and which is not? (Select one)">
    A. Exact string match against a single reference answer.
    B. A rubric with an LLM-as-judge in a separate session (and pairwise A/B for comparisons), not exact match.
    C. Aggregate accuracy only.
    D. The generator grading itself for speed.

    **Answer: B.** Open-ended output has no single correct string, so use a rubric + separate-session judge, with pairwise A/B for comparisons. Exact match (A) fits deterministic labels, not prose; aggregate-only (C) masks segments; self-grading (D) is anti-pattern #9.
  </AccordionItem>

  <AccordionItem title="Q14 · A production call intermittently returns `400 invalid_request` and is retried with exponential backoff, never succeeding. What is the correct triage? (Select one)">
    A. Add more retries with longer backoff.
    B. Recognise a `400` is a deterministic integration/schema error (e.g. a forbidden parameter such as forced `tool_choice` on Fable 5.1); fix the request shape rather than retrying.
    C. Treat it as a model hallucination and rewrite the prompt.
    D. Rotate the API key.

    **Answer: B.** A `400` will never succeed on retry; read it and fix the request. More backoff (A) loops forever; it is not a model-output problem (C); auth rotation (D) addresses `401`/`403`, not a `400`.
  </AccordionItem>

  <AccordionItem title="Q15 · A trace shows an agent repeating the same tool call with growing input tokens, ending in `stop_reason: max_tokens`. What is the root cause? (Select one)">
    A. The model ran out of output tokens; raise `max_tokens`.
    B. A loop-control defect: the harness is not feeding the tool result back and/or termination is not driven by `stop_reason`; fix the loop. The final `max_tokens` is a downstream symptom.
    C. Rate limiting; add backoff.
    D. A hallucination; add citations.

    **Answer: B.** Repeating the same call with growing context is a loop-control problem (anti-patterns #1/#2); the `max_tokens` end is a symptom, not the cause (A). It is not transport (C) or grounding (D).
  </AccordionItem>

  <AccordionItem title="Q16 · A CI eval gates only on aggregate accuracy (93%), and a handwritten-forms regression ships. What TWO changes fix the eval process? (Select two)">
    A. Report and gate on per-segment metrics (by document type), failing if any segment drops below its threshold.
    B. Use an LLM-as-judge in a separate session/model against a rubric for the open-ended parts.
    C. Raise the aggregate threshold to 96%.
    D. Have the same session grade its own output.
    E. Run evals only monthly.

    **Answer: A and B.** Per-segment gating surfaces the hidden failure (anti-pattern #10) and a separate-session judge avoids self-review bias (anti-pattern #9). A higher aggregate (C) still masks segments; self-grading (D) is #9; monthly evals (E) slow detection.
  </AccordionItem>

  <AccordionItem title="Q17 · A helper returns an empty list when the Claude call throws, and the caller treats empty as 'no matches'. Which anti-pattern is this and the fix? (Select one)">
    A. Aggregate metrics; add per-segment reporting.
    B. Silent error suppression (anti-pattern #7); propagate the error with diagnostic context instead of returning empty-as-success.
    C. Same-session self-review; use a separate judge.
    D. Iteration cap; drive from `stop_reason`.

    **Answer: B.** Returning empty on failure reports failure as success — silent suppression (anti-pattern #7). Surface the error with context. The others name unrelated anti-patterns.
  </AccordionItem>

  <AccordionItem title="Q18 · When comparing two prompt versions with an LLM judge, results flip depending on which answer is shown first. What is the issue and the fix? (Select one)">
    A. The judge model is broken; switch models.
    B. Position bias in pairwise judging; randomise the A/B order (and optionally average both orders) so position does not decide the winner.
    C. Temperature is too low; raise it.
    D. Use exact match instead.

    **Answer: B.** Pairwise judges can be swayed by answer position; randomising order (or scoring both orders) removes the bias. It is not a broken model (A); temperature (C) is not the cause; exact match (D) does not fit open-ended comparison.
  </AccordionItem>
</Accordions>

## Key takeaways

- Triage by layer: integration errors (429/5xx/timeouts/JSON/roles) get retries and request fixes; model-output errors get prompt/context/model changes and validation.
- Log request IDs, model, `stop_reason`, `usage` and latency for trace analysis.
- Reproduce with a pinned snapshot and `temperature: 0`, accepting residual non-determinism.
- Evaluate with a golden set and rubric; use LLM-as-judge in a *separate* session (never same-session self-review).
- Run regression tests in CI and report per-segment metrics – aggregates hide the failures that matter.
- Match the eval method to the output: exact match for labels, field-level match for extraction, rubric + separate-session LLM-judge for prose, pairwise A/B (with randomised order) for comparisons.
- Triage with the status code and `stop_reason` first: `429`/timeout/JSON = integration; hallucination/off-schema/refusal = model output — the wrong-layer fix is the classic distractor.
- Gate CI on the worst-performing segment, not the aggregate; raising the aggregate threshold still masks a failing segment.
- A deterministic `400` never succeeds on retry — fix the request shape (e.g. forced tool choice on Fable 5.1) rather than adding backoff.
