# Evals Cookbook

Golden-set design, graders, LLM-as-judge calibration, A/B testing with significance, CI regression YAML, per-segment reporting and cost/latency dashboards.

import { Steps } from '@prosefly/astro-components';

Evals turn "it seems better" into evidence. The exam-correct posture: a **golden set**, the **cheapest sufficient grader**, judges in a **separate session**, **per-segment** reporting, and a **CI regression gate** on prompt/model changes.

:::tip[Exam signal]
"Overall accuracy is 95%" is a trap — demand per-segment metrics. "Are you sure?" in the same chat is a trap — use a fresh session/model. Comparing two prompts without significance is a trap — use pairwise A/B with a significance test.
:::

## Golden-set design

<Steps>
1. **Cover the distribution** — sample real inputs across segments (document type, language, difficulty, edge cases), not just easy ones.
2. **Include hard/adversarial cases** — injection attempts, ambiguous inputs, known past failures.
3. **Label with expected outputs or rubrics** — exact answers where possible; rubrics for subjective tasks.
4. **Size** — enough per segment to detect regressions (aim ≥ 30–50 per segment for stable rates).
5. **Version and freeze** — a golden set is a controlled asset; changing it invalidates comparisons.
6. **Separate holdout** — keep a set you never tune against to catch overfitting to the eval.
</Steps>

## Graders — cheapest sufficient first

| Grader | Use when | Cost | Reliability |
| --- | --- | --- | --- |
| Exact / normalised match | Deterministic answers (classification, extraction fields) | Free | High |
| Regex / structural | Format, presence of fields, schema validity | Free | High |
| Heuristic (numeric tolerance, set overlap) | Numbers, lists | Free | Medium |
| Rubric grading | Subjective quality with clear criteria | Low | Medium |
| LLM-as-judge | Open-ended quality, pairwise preference | Higher | Medium (needs calibration) |
| Human | Ground truth, calibration reference | Highest | Highest |

Escalate only as far as the task requires — exact match before rubric before LLM-judge.

## LLM-as-judge calibration

The judge must run in a **separate session, ideally a different model**, and be **calibrated against human labels** before you trust it.

<Steps>
1. Draft the rubric with explicit, non-overlapping criteria and a scale.
2. Have humans label a calibration set (e.g., 100 items).
3. Run the judge on the same set; measure agreement (e.g., Cohen's κ, or % agreement).
4. If agreement is low, tighten the rubric, add few-shot exemplars of each grade, or reduce the scale (binary is easier to calibrate than 1–10).
5. Re-measure; only deploy the judge once agreement meets your bar (e.g., κ ≥ 0.6).
6. Re-calibrate when the model, rubric or task shifts.
</Steps>

| Calibration symptom | Fix |
| --- | --- |
| Judge too lenient | Add negative exemplars; sharpen "fail" criteria |
| Judge inconsistent run-to-run | Lower temperature; binary rubric; majority vote of 3 |
| Position bias in pairwise | Randomise A/B order; average both orderings |
| Verbosity bias | Instruct to ignore length; penalise unsupported claims |

## A/B testing with significance

Pairwise comparison of prompt/model B against A. Do not eyeball a few outputs — test.

Worked example: B wins 118 of 200 head-to-head comparisons (ties split). Is B really better?

```text
n = 200, wins = 118, p̂ = 0.59
Two-sided test of p = 0.5:
  z = (118 − 100) / sqrt(200 × 0.5 × 0.5) = 18 / 7.07 ≈ 2.55
  p-value ≈ 0.011  → significant at α = 0.05
95% CI on win rate: 0.59 ± 1.96 × sqrt(0.59×0.41/200) = 0.59 ± 0.068 → [0.52, 0.66]
```

B is significantly better. Had B won 108/200: `z ≈ 1.13`, p ≈ 0.26 — **not** significant; do not ship on that alone.

| Pitfall | Fix |
| --- | --- |
| Too few samples | Power the test; more items for small effects |
| Peeking / stopping early | Fix n in advance or use sequential-test corrections |
| Ignoring ties | Define handling up front |
| Aggregate only | Also test per segment |

## CI regression gate

Run the golden set on every prompt/model change; block the deploy if a segment regresses beyond tolerance.

```yaml
# .github/workflows/evals.yml
name: evals
on: [pull_request]
jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: pip install -r eval/requirements.txt
      - name: Run golden set
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
        run: python eval/run.py --golden eval/golden.jsonl --out eval/report.json
      - name: Gate
        run: |
          python - <<'PY'
          import json, sys
          r = json.load(open("eval/report.json"))
          # per-segment gate: no segment below its baseline minus 2 points
          fails = [s for s, m in r["by_segment"].items() if m["accuracy"] < r["baseline"][s] - 0.02]
          if fails:
              print("Regressed segments:", fails); sys.exit(1)
          print("All segments within tolerance.")
          PY
```

Gate on **per-segment** thresholds, not just the aggregate — that is what catches a single failing document type.

## Per-segment reporting

```text
Segment          n     Accuracy   Faithfulness   Δ vs baseline
--------------   ---   --------   ------------   -------------
invoices         120   0.94       0.95            +0.01
contracts         90   0.88       0.91            −0.00
receipts (scan)   60   0.61 ◄     0.72 ◄          −0.18 ◄
--------------   ---   --------   ------------   -------------
OVERALL          270   0.86       0.90            −0.03
```

Aggregate 0.86 looks fine; the scanned-receipts segment at 0.61 is the real story. Always report the breakdown.

## Cost / latency dashboards

| Metric | Why | Watch for |
| --- | --- | --- |
| Cost per request / per 1k | Budget | Spike after a model or prompt change |
| Tokens in/out (p50/p95) | Cost + context pressure | Growing prompts → cache or trim |
| Latency p50 / p95 / p99 | UX / SLO | p95 breach even when p50 is fine |
| Cache hit rate | Cost efficiency | Drop → prefix churned |
| Error rate by type (429/5xx/refusal) | Reliability | 429 → rate-limit tier; refusals → prompt/policy |
| Fallback rate | Capacity | High → capacity or model issue |
| Quality (eval score) over time | Regression | Silent drift after upstream changes |

Report **p95/p99**, not just averages — an average hides the tail that violates the SLO.

## Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "95% overall means we're good" | A segment may fail; report per-segment | Aggregate-metric anti-pattern |
| "The model can grade its own answer" | Use a separate session/model, calibrated | Same-session-review anti-pattern |
| "B looked better in a few outputs" | Test significance with enough samples | Eyeballing distractor |
| "LLM-judge is objective" | Needs calibration vs humans; has biases | Uncalibrated-judge distractor |
| "Average latency meets the SLO" | Watch p95/p99 tails | Aggregate-latency distractor |
| "Evals are a one-time thing" | CI regression + online monitoring, continuous | Point-in-time distractor |
| "Change the golden set to pass" | Freeze it; changing invalidates comparison | Overfitting distractor |

## Scenario walkthrough

A team wants to switch the extraction pipeline from Sonnet 5 to Haiku 4.5 to cut cost. Aggregate accuracy on a quick test looks "about the same". How do you decide correctly?

<Steps>
1. **Golden set, per-segment** — run both models on the frozen set, break results down by document type.
2. **Right grader** — exact/structural match on the extracted fields (deterministic), not an LLM judge.
3. **Significance** — where subjective, pairwise A/B with a significance test, not eyeballing.
4. **Find the hidden regression** — Haiku matches on clean PDFs but drops 18 points on scanned receipts (aggregate hid it).
5. **Decide by constraint** — if scanned receipts matter, keep Sonnet 5 for that segment (cascade), Haiku for the rest.
6. **Gate it** — add the per-segment thresholds to CI so a future switch cannot silently regress.
7. **Watch cost + latency** — confirm the saving is real on the dashboard (p95, cost per 1k).
</Steps>

Rejected alternatives: switching on aggregate accuracy (aggregate-metric), judging with the same model in the same session (same-session review), and eyeballing "about the same" without significance (eyeballing distractor).

## Key takeaways

- Golden set first: cover the distribution, include hard cases, version and freeze it.
- Use the **cheapest sufficient grader**; LLM-as-judge must be separate-session and calibrated.
- Prove improvements with **pairwise A/B + significance**, not a few outputs.
- Gate deploys with a **CI regression** suite on **per-segment** thresholds.
- Monitor cost and latency at **p95/p99**, and quality continuously — evals are ongoing, not one-time.
