AI Cert Prep
Type to search documentation.

Appendix · OpenAI

Evals Cookbook (OpenAI)

Why evals precede prompt tuning, building datasets from real traffic and failure logs, grader types and when each is valid, offline versus online evaluation, CI regression gates, the failure-analysis loop, agent-workflow evaluation, the prompt optimizer, where the Evals API sits in the docs, and a worked eval spec.

Evals turn “it seems better” into evidence. This cookbook mirrors the objectives of the Academy Evaluate AI Applications course; it is independent preparation. The correct posture it teaches: build a dataset from real traffic and failures, pick the cheapest sufficient grader, run graders that are independent of the thing under test, report per segment, and gate changes with a CI regression suite before any online rollout.

Assessment signal

“Overall accuracy is 95%” is a trap — demand the per-segment breakdown. “We changed the prompt and it feels better” is a trap — where is the eval and the baseline? Building the golden set by paraphrasing the docs instead of sampling real traffic is a trap — the eval then tests the docs, not your users.

Why evals precede prompt tuning

Prompt tuning without an eval is guessing with confidence. You cannot tell whether a change helped, hurt, or helped one segment while breaking another. The sequence the Evaluate AI Applications objectives insist on:

text
define the task and success criteria
│
▼
build a dataset (real traffic + failures) ◄── the eval is the ruler
│
▼
measure the current baseline
│
▼
change ONE thing (prompt / model / retrieval)
│
▼
re-measure ─── better? keep. worse/flat? revert.

Build the ruler before you start cutting. A prompt change that “feels better” on three examples routinely regresses a segment you were not looking at; only a baseline plus a dataset can catch that.

Dataset construction from real traffic and failure logs

  1. Sample from production, across segments — document type, language, difficulty, user cohort — not just easy or synthetic cases.
  2. Mine the failure logs — every reported wrong answer, refusal, or complaint is a labelled hard case waiting to be added; failures are the highest-yield source.
  3. Label with expected outputs or rubrics — exact answers where the task is deterministic; a clear rubric where it is subjective.
  4. Size for stable rates — aim for enough items per segment (roughly 30–50) to detect a regression rather than noise.
  5. Version and freeze — the dataset is a controlled asset; changing it invalidates comparisons across runs.
  6. Keep a holdout — a set you never tune against, so you can catch overfitting to the eval itself.

Sanitise real traffic before it enters the dataset: strip or mask PII, and store only what the eval needs. A golden set of real cases is far more predictive than one paraphrased from marketing copy.

Grader types and when each is valid

GraderValid whenCostReliability
Exact / normalised matchDeterministic answers: classification labels, extracted fieldsFreeHigh
Code grader (assertions, regex, structural)Format, schema validity, numeric tolerance, set overlapFreeHigh
Rubric / model grader (model-as-judge)Open-ended quality with clear, non-overlapping criteriaHigherMedium — must be calibrated
HumanGround truth and calibration referenceHighestHighest

Escalate only as far as the task requires: try exact match, then a code grader, then a model grader, then human. A model grader must run independently of the model under test — a separate request, ideally a different model — and must be calibrated against human labels before you trust it. The calibration loop: humans label a set, the model grader scores the same set, you measure agreement, tighten the rubric or add graded exemplars until agreement clears your bar, then re-calibrate whenever the task, rubric or model shifts.

Model-grader symptomFix
Too lenientAdd negative exemplars; sharpen the “fail” criteria
Inconsistent run-to-runLower temperature; use a binary rubric; take the majority of three
Position bias in pairwiseRandomise A/B order; average both orderings
Verbosity biasInstruct it to ignore length; penalise unsupported claims

Offline vs online evaluation

OfflineOnline
WhenBefore shipping a changeAfter shipping, on live traffic
DataFrozen golden setReal user interactions
CatchesRegressions against known casesDistribution shift, novel failures
RiskCannot see what the set omitsReal users experience failures first
ToolingEval run in CILogging, sampling, monitors, human review of sampled traffic

They are complementary, not alternatives. Offline evals gate the change; online evaluation catches what the golden set never contained — new query shapes, drift, and edge cases that then feed back into the offline set. Shipping on offline evals alone assumes your golden set already knows every failure, which it does not.

Regression gates in CI

Run the golden set on every prompt, model or retrieval change and block the merge if a segment regresses beyond tolerance.

.github/workflows/evals.yml
name: evals
on: [pull_request]
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: pip install -r eval/requirements.txt
- name: Run golden set
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
run: python eval/run.py --golden eval/golden.jsonl --out eval/report.json
- name: Gate on per-segment thresholds
run: python eval/gate.py --report eval/report.json --tolerance 0.02

Gate on per-segment thresholds, not just the aggregate — that is what catches a single failing document type or language before it reaches users.

The failure-analysis loop

  1. Collect — pull the failing cases from evals and from production logs.
  2. Cluster — group by root cause: retrieval miss, prompt ambiguity, missing context, model limitation, bad label.
  3. Attribute — assign each cluster to the stage that owns it, so the fix lands in the right place.
  4. Fix the largest cluster first — the biggest cluster is the highest-leverage change.
  5. Add the cases to the golden set — a fixed failure that is not in the eval will silently regress later.
  6. Re-measure and repeat — the loop is continuous, not a one-time audit.

Agent-workflow evaluation

Single-turn accuracy does not tell you whether an agent completed the job. Evaluate the trajectory as well as the answer.

DimensionWhat to check
Task successDid the workflow reach the correct end state?
Tool-call correctnessRight tools, right arguments, right order
Step efficiencyDid it reach the goal without wasteful or looping steps?
RecoveryDid it handle a tool error or empty result gracefully?
Safety / boundariesDid it stay within its stated boundaries (no out-of-scope or irreversible actions)?
Cost and latencyTokens, tool calls, wall-clock per completed task

An agent that produces a plausible final message while calling the wrong tools has failed — trajectory evaluation catches what an answer-only grader misses.

The prompt optimizer

The docs provide a prompt optimizer that refines a prompt against a dataset and graders rather than by intuition. Use it as a disciplined loop, not a magic wand: supply a representative dataset and a valid grader, let it propose prompt variants, and keep a variant only if it beats the baseline on the frozen set per segment. It automates the search over prompt phrasings; it does not replace the dataset, the grader, or your judgement about which metric matters.

Where the Evals API sits in the docs

In September 2026 the docs group Evals and fine-tuning under “Legacy APIs” alongside Agent Builder and the Assistants API. This is a documentation-organisation fact, not a verdict on the practice: evaluation remains correct, essential engineering. Say “the Evals API (now grouped with the legacy APIs)” rather than presenting it as the newest surface, and do not infer that evals are deprecated as a discipline — they are not. Fine-tuning likewise remains available (supervised, vision, DPO, and reinforcement fine-tuning) for style and behaviour, not for injecting facts.

A worked eval spec

A concrete spec for an invoice-extraction pipeline you could run tomorrow.

yaml
eval: invoice-extraction-v3
task: Extract {invoice_number, total, currency, due_date} from an invoice document.
dataset:
source: 300 real invoices sampled across 4 vendors + 40 past-failure cases
segments: [clean_pdf, scanned, non_english, multi_page]
labels: gold field values, hand-checked; PII masked
frozen: true
holdout: 60 cases never used for tuning
graders:
- fields: exact match on invoice_number, currency
kind: code
- total: numeric within 0.005 tolerance
kind: code
- due_date: parses to ISO date and equals gold
kind: code
baseline:
model: gpt-5.6-terra
by_segment: { clean_pdf: 0.97, scanned: 0.79, non_english: 0.88, multi_page: 0.91 }
gate:
rule: no segment below its baseline minus 0.02
online:
sample: 2% of production, human-review weekly, feed failures back into dataset

Reading the spec: exact and numeric code graders because the task is deterministic (no model-judge needed or wanted); a scanned segment already weakest at 0.79, so a model switch that lifts the aggregate but drops scanned further must be blocked; a holdout to catch overfitting; and an online sample so novel failures rejoin the dataset. Report per segment always — an aggregate of 0.91 here hides the scanned segment that users actually complain about.

Common misconceptions

MisconceptionRealityWhy it matters on the assessment
“95% overall means we are good”A segment may be failing; report per segmentAggregate-metric trap
“Tune the prompt, then maybe add an eval”The eval is the ruler; build it firstSequence trap
“The model can grade its own output in the same thread”Grade in a separate request, calibratedIndependence trap
“A model grader is objective”It needs calibration against humans and has biasesUncalibrated-judge distractor
“Offline evals are enough to ship”Online evaluation catches distribution shiftOffline-only trap
“Evals are grouped as legacy, so skip them”Legacy in the docs’ layout, still correct practiceDocs-organisation trap
“Paraphrase the docs to build the golden set”Sample real traffic and failuresDataset-source trap

Scenario walkthrough

A team wants to switch an extraction pipeline from gpt-5.6-terra to gpt-5.6-luna to cut cost. A quick test shows aggregate accuracy “about the same”. How do you decide correctly?

  1. Golden set, per segment — run both models on the frozen dataset and break the results down by document type.
  2. Right grader — exact and numeric code graders on the extracted fields; this is deterministic, so a model judge would only add noise.
  3. Find the hidden regression — Luna matches Terra on clean PDFs but drops sharply on scanned invoices; the aggregate hid it.
  4. Decide by constraint — if scanned invoices matter, keep Terra for that segment and route the rest to Luna, capturing the real saving without the regression.
  5. Gate it — add the per-segment thresholds to CI so a future switch cannot silently regress.
  6. Watch online — sample production, confirm the cost saving is real, and feed any new failures back into the dataset.

Rejected alternatives: switching on aggregate accuracy (aggregate-metric trap), judging deterministic fields with a model grader in the same session (independence trap), and eyeballing “about the same” with no baseline or significance (no-ruler trap).

Key takeaways

  • Build the eval before you tune: the dataset is the ruler, drawn from real traffic and failure logs, versioned and frozen.
  • Use the cheapest sufficient grader; model graders must be independent and calibrated against human labels.
  • Combine offline evals (gate the change in CI, per segment) with online evaluation (catch drift and novel failures).
  • Run a continuous failure-analysis loop and feed every fixed failure back into the golden set.
  • Evaluate agent trajectories, not just final answers: tool correctness, efficiency, recovery, boundaries.
  • Evals and fine-tuning are grouped under the docs’ legacy APIs but remain correct, essential practice.

Last updated Sep 18, 2026