Appendix · Claude
Evals Cookbook
Golden-set design, graders, LLM-as-judge calibration, A/B testing with significance, CI regression YAML, per-segment reporting and cost/latency dashboards.
Evals turn “it seems better” into evidence. The exam-correct posture: a golden set, the cheapest sufficient grader, judges in a separate session, per-segment reporting, and a CI regression gate on prompt/model changes.
Exam signal
“Overall accuracy is 95%” is a trap — demand per-segment metrics. “Are you sure?” in the same chat is a trap — use a fresh session/model. Comparing two prompts without significance is a trap — use pairwise A/B with a significance test.
Golden-set design
- Cover the distribution — sample real inputs across segments (document type, language, difficulty, edge cases), not just easy ones.
- Include hard/adversarial cases — injection attempts, ambiguous inputs, known past failures.
- Label with expected outputs or rubrics — exact answers where possible; rubrics for subjective tasks.
- Size — enough per segment to detect regressions (aim ≥ 30–50 per segment for stable rates).
- Version and freeze — a golden set is a controlled asset; changing it invalidates comparisons.
- Separate holdout — keep a set you never tune against to catch overfitting to the eval.
Graders — cheapest sufficient first
| Grader | Use when | Cost | Reliability |
|---|---|---|---|
| Exact / normalised match | Deterministic answers (classification, extraction fields) | Free | High |
| Regex / structural | Format, presence of fields, schema validity | Free | High |
| Heuristic (numeric tolerance, set overlap) | Numbers, lists | Free | Medium |
| Rubric grading | Subjective quality with clear criteria | Low | Medium |
| LLM-as-judge | Open-ended quality, pairwise preference | Higher | Medium (needs calibration) |
| Human | Ground truth, calibration reference | Highest | Highest |
Escalate only as far as the task requires — exact match before rubric before LLM-judge.
LLM-as-judge calibration
The judge must run in a separate session, ideally a different model, and be calibrated against human labels before you trust it.
- Draft the rubric with explicit, non-overlapping criteria and a scale.
- Have humans label a calibration set (e.g., 100 items).
- Run the judge on the same set; measure agreement (e.g., Cohen’s κ, or % agreement).
- If agreement is low, tighten the rubric, add few-shot exemplars of each grade, or reduce the scale (binary is easier to calibrate than 1–10).
- Re-measure; only deploy the judge once agreement meets your bar (e.g., κ ≥ 0.6).
- Re-calibrate when the model, rubric or task shifts.
| Calibration symptom | Fix |
|---|---|
| Judge too lenient | Add negative exemplars; sharpen “fail” criteria |
| Judge inconsistent run-to-run | Lower temperature; binary rubric; majority vote of 3 |
| Position bias in pairwise | Randomise A/B order; average both orderings |
| Verbosity bias | Instruct to ignore length; penalise unsupported claims |
A/B testing with significance
Pairwise comparison of prompt/model B against A. Do not eyeball a few outputs — test.
Worked example: B wins 118 of 200 head-to-head comparisons (ties split). Is B really better?
n = 200, wins = 118, p̂ = 0.59Two-sided test of p = 0.5: z = (118 − 100) / sqrt(200 × 0.5 × 0.5) = 18 / 7.07 ≈ 2.55 p-value ≈ 0.011 → significant at α = 0.0595% CI on win rate: 0.59 ± 1.96 × sqrt(0.59×0.41/200) = 0.59 ± 0.068 → [0.52, 0.66]B is significantly better. Had B won 108/200: z ≈ 1.13, p ≈ 0.26 — not significant; do not ship on that alone.
| Pitfall | Fix |
|---|---|
| Too few samples | Power the test; more items for small effects |
| Peeking / stopping early | Fix n in advance or use sequential-test corrections |
| Ignoring ties | Define handling up front |
| Aggregate only | Also test per segment |
CI regression gate
Run the golden set on every prompt/model change; block the deploy if a segment regresses beyond tolerance.
name: evalson: [pull_request]jobs: eval: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - run: pip install -r eval/requirements.txt - name: Run golden set env: ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} run: python eval/run.py --golden eval/golden.jsonl --out eval/report.json - name: Gate run: | python - <<'PY' import json, sys r = json.load(open("eval/report.json")) # per-segment gate: no segment below its baseline minus 2 points fails = [s for s, m in r["by_segment"].items() if m["accuracy"] < r["baseline"][s] - 0.02] if fails: print("Regressed segments:", fails); sys.exit(1) print("All segments within tolerance.") PYGate on per-segment thresholds, not just the aggregate — that is what catches a single failing document type.
Per-segment reporting
Segment n Accuracy Faithfulness Δ vs baseline-------------- --- -------- ------------ -------------invoices 120 0.94 0.95 +0.01contracts 90 0.88 0.91 −0.00receipts (scan) 60 0.61 ◄ 0.72 ◄ −0.18 ◄-------------- --- -------- ------------ -------------OVERALL 270 0.86 0.90 −0.03Aggregate 0.86 looks fine; the scanned-receipts segment at 0.61 is the real story. Always report the breakdown.
Cost / latency dashboards
| Metric | Why | Watch for |
|---|---|---|
| Cost per request / per 1k | Budget | Spike after a model or prompt change |
| Tokens in/out (p50/p95) | Cost + context pressure | Growing prompts → cache or trim |
| Latency p50 / p95 / p99 | UX / SLO | p95 breach even when p50 is fine |
| Cache hit rate | Cost efficiency | Drop → prefix churned |
| Error rate by type (429/5xx/refusal) | Reliability | 429 → rate-limit tier; refusals → prompt/policy |
| Fallback rate | Capacity | High → capacity or model issue |
| Quality (eval score) over time | Regression | Silent drift after upstream changes |
Report p95/p99, not just averages — an average hides the tail that violates the SLO.
Common misconceptions
| Misconception | Reality | Why it matters on the exam |
|---|---|---|
| “95% overall means we’re good” | A segment may fail; report per-segment | Aggregate-metric anti-pattern |
| “The model can grade its own answer” | Use a separate session/model, calibrated | Same-session-review anti-pattern |
| “B looked better in a few outputs” | Test significance with enough samples | Eyeballing distractor |
| “LLM-judge is objective” | Needs calibration vs humans; has biases | Uncalibrated-judge distractor |
| “Average latency meets the SLO” | Watch p95/p99 tails | Aggregate-latency distractor |
| “Evals are a one-time thing” | CI regression + online monitoring, continuous | Point-in-time distractor |
| “Change the golden set to pass” | Freeze it; changing invalidates comparison | Overfitting distractor |
Scenario walkthrough
A team wants to switch the extraction pipeline from Sonnet 5 to Haiku 4.5 to cut cost. Aggregate accuracy on a quick test looks “about the same”. How do you decide correctly?
- Golden set, per-segment — run both models on the frozen set, break results down by document type.
- Right grader — exact/structural match on the extracted fields (deterministic), not an LLM judge.
- Significance — where subjective, pairwise A/B with a significance test, not eyeballing.
- Find the hidden regression — Haiku matches on clean PDFs but drops 18 points on scanned receipts (aggregate hid it).
- Decide by constraint — if scanned receipts matter, keep Sonnet 5 for that segment (cascade), Haiku for the rest.
- Gate it — add the per-segment thresholds to CI so a future switch cannot silently regress.
- Watch cost + latency — confirm the saving is real on the dashboard (p95, cost per 1k).
Rejected alternatives: switching on aggregate accuracy (aggregate-metric), judging with the same model in the same session (same-session review), and eyeballing “about the same” without significance (eyeballing distractor).
Key takeaways
- Golden set first: cover the distribution, include hard cases, version and freeze it.
- Use the cheapest sufficient grader; LLM-as-judge must be separate-session and calibrated.
- Prove improvements with pairwise A/B + significance, not a few outputs.
- Gate deploys with a CI regression suite on per-segment thresholds.
- Monitor cost and latency at p95/p99, and quality continuously — evals are ongoing, not one-time.
Last updated Sep 18, 2026