AI Cert Prep
Type to search documentation.

Domains

D4 · Evaluation, Testing & Optimization

Evaluation metrics, building eval datasets, harness design, rubric grading and LLM-as-judge, A/B testing and significance, offline vs online evals, regression suites, failure diagnosis, and the cost/latency optimisation playbook.

This domain is worth 16% – roughly 10 of 63 items. It tests whether you can prove a system works and keep it working: choosing the right metrics, building trustworthy eval datasets, running calibrated LLM-as-judge and A/B tests, catching regressions in CI, diagnosing which layer failed, and optimising cost and latency without degrading quality. The dominant judgement: measure per-segment with independent evaluation, never aggregate-only or self-report.

Learning objectives

By the end of this page you should be able to:

  1. Choose evaluation metrics (accuracy, latency p50/p95, cost per task, safety, security) per use case.
  2. Build eval datasets: golden sets, synthetic data, edge cases, per-segment coverage.
  3. Design a test harness and rubric grading; run LLM-as-judge in a separate session/model and calibrate it.
  4. Run pairwise A/B tests and reason about statistical significance.
  5. Distinguish offline vs online evals and run regression suites in CI.
  6. Diagnose prompt failure vs hallucination vs model mismatch vs retrieval failure.
  7. Apply the cost/latency/token optimisation playbook and roll out changes via canary.

4.1 Evaluation metrics

Different use cases have different definitions of “good”. An architect picks the metrics that map to the success criteria (D1).

MetricDefinitionWhen it dominates
Accuracy / task successCorrect outcome vs ground truth or rubricMost quality-critical tasks
Latency p50 / p95Median / tail response timeInteractive/SLA-bound systems (watch p95, not just p50)
Cost per task$ per completed unit of workHigh-volume, cost-sensitive systems
SafetyRate of harmful/policy-violating outputsUser-facing, regulated
SecurityPrompt-injection/exfiltration resistanceTool-using agents, untrusted inputs

Exam signal

“Average latency is fine but some users wait 12 s” → you are being tested on p95/tail latency, not the mean. “Overall accuracy is 92%” while one segment fails → the answer is per-segment metrics, not the aggregate (anti-pattern #10).


4.2 Building eval datasets

An eval is only as trustworthy as its dataset. Coverage matters more than size.

Dataset elementPurposeSourcing
Golden setCurated inputs with verified expected outputsHand-labelled by SMEs; the source of truth
Synthetic dataScale coverage cheaply; explore rare inputsGenerate with a model, then human-review a sample
Edge casesExercise the failure boundaryMined from production errors, adversarial inputs
Per-segment coverageEnsure every document type / customer tier / language is representedStratify the set so no segment is invisible

Aggregate blindness

A golden set that is 90% one document type will report high accuracy while the 10% type silently fails. Stratify the dataset and report metrics per segment.


4.3 Harness design and rubric grading

A test harness runs each eval input through the system and scores the output. Scoring methods, from cheapest to richest:

MethodUse whenCaveat
Exact / string / regex matchDeterministic outputs (IDs, classifications, JSON fields)Brittle for free text
Rubric gradingStructured criteria (e.g. correctness, completeness, tone each 0–2)Needs a clear rubric; use LLM-as-judge or humans to apply it
LLM-as-judgeFree-text quality at scaleMust use a separate model/session and be calibrated
Human reviewHighest-stakes, ambiguous, or calibration referenceSlow/expensive; use as the gold standard

A rubric turns fuzzy quality into checkable dimensions:

text
Criterion Weight Scale
Correctness 0.5 0 wrong · 1 partial · 2 fully correct
Grounding 0.3 0 unsupported · 1 partial · 2 fully cited
Tone/format 0.2 0 off · 1 minor · 2 on-spec
Score = Σ (weight × normalised criterion)

4.4 LLM-as-judge and calibration

LLM-as-judge scales evaluation, but only if it is trustworthy.

  • Use a different model or at least a fresh session than the one under test — same-session self-review carries the original reasoning bias (anti-pattern #9).
  • Calibrate the judge against a set of human labels: measure agreement (e.g. accuracy vs human, correlation). If the judge disagrees with humans, fix the rubric/prompt before trusting it at scale.
  • Watch for judge biases: position bias (favours the first option), length bias (favours longer answers), self-preference (favours its own family’s style).
text
human-labelled sample ──▶ run judge ──▶ compare
│ │
└────── agreement ≥ threshold? ──── yes ─▶ trust judge at scale
no ─▶ revise rubric/judge prompt, recalibrate

Never let the model grade its own homework in-session

Asking the producing session “is this correct?” retains the bias that produced the answer. Evaluation must be independent.


4.5 Pairwise A/B testing and statistical significance

To compare prompt A vs prompt B (or model A vs B), pairwise comparison on the same inputs is more sensitive than comparing separate averages.

  1. Run both variants on the same eval inputs (paired design).

  2. Have an independent judge/human pick the winner per item (or score each).

  3. Compute the win rate and test whether the difference is statistically significant — a 52% win on 30 items is noise; the same on 2,000 items may be real. Use a significance test and report the confidence interval.

  4. Only promote the winner if the improvement is significant and holds per segment (no regression on any segment).

Exam signal

“Prompt B looks a bit better on a handful of examples” → the correct answer invokes statistical significance / larger sample, not “ship B because it looks better”. Small-sample eyeballing is a trap.


4.6 Offline vs online evaluation

Offline evalOnline eval
WhenPre-release, in CIIn production, live traffic
DataGolden/synthetic setsReal user interactions
MeasuresCorrectness vs known answers, regressionsReal outcomes: deflection, thumbs, task completion, cost
RiskMay not reflect production distributionExposes users to changes (mitigate with canary)

You need both: offline gates releases; online catches distribution shift the golden set missed.

Regression suites in CI

Every prompt, model or retrieval change must run the regression suite before promotion. A regression suite is the accumulated golden set plus every past production failure turned into a test. This is how you prevent a fix for one segment from breaking another.

text
PR opened ─▶ CI runs regression suite (offline evals, per-segment)
├─ any segment regresses ─▶ block merge
└─ all pass ─▶ allow ─▶ canary online ─▶ ramp

4.7 Diagnosing failures: which layer broke?

When output is wrong, isolate the layer before “fixing” anything. Fixing the wrong layer is the most expensive mistake in this domain.

SymptomLikely causeConfirm byFix at
Answer uses facts not in provided contextHallucination / weak groundingCheck faithfulness against retrieved chunksGrounding instruction, citations
Right chunks retrieved, wrong answer shape/formatPrompt failureInspect prompt vs output; test prompt in isolationPrompt/template
Confident-but-wrong after a data changeRetrieval/indexing (stale)Log retrieved chunk IDsRe-index / freshness pipeline (D3)
Fails only on hardest cases, fine elsewhereModel mismatch (too small)Re-run failures on a stronger modelRoute/escalate model tier
Fails only on one segmentCoverage / segment-specific issuePer-segment metricsTargeted data/prompt/retrieval fix
text
Wrong output ─▶ Was the right context retrieved?
├─ No ─▶ retrieval/indexing failure (D3)
└─ Yes ─▶ Is the answer unsupported by that context?
├─ Yes ─▶ hallucination / grounding
└─ No ─▶ Is only the hardest tier failing?
├─ Yes ─▶ model mismatch (escalate tier)
└─ No ─▶ prompt/format failure

4.8 The cost / latency / token optimisation playbook

Optimise only after you can measure (per-segment) and without regressing the quality bar.

LeverCutsTrade-off / caveat
Prompt caching (stable-prefix-first)Input cost, latencyNeeds stable prefix; ~1024-token minimum
Batching (Message Batches API)50% costUp to 24 h latency — only latency-tolerant work
Routing / cascadesCostAdds a classifier/validation step
Output trimming / structured outputsOutput cost, latencyDon’t drop needed content
Effort tuning (lower where adequate)Thinking tokens, latencyToo low hurts hard tasks
Model right-sizingCost, latencyRe-validate quality on the smaller model
Fewer tool round-trips / progressive discoveryLatency, tokensRequires tool-set discipline
text
Cost too high? ─▶ cache stable prefix ─▶ route cheap-first ─▶ batch tolerant work ─▶ trim output
Latency too high? ─▶ smaller/faster model ─▶ fast mode ─▶ fewer round-trips ─▶ lower effort ─▶ stream
(after each change: re-run the regression suite per segment)

Exam signal

Optimisation answers that skip re-validation are wrong. Every cost/latency change must be re-evaluated per segment — a cheaper model or lower effort that quietly regresses one segment is a failure, not a win.


4.9 Canary rollouts of prompts and models

Never flip a prompt or model change to 100% at once.

  1. Pass the offline regression suite (per segment).

  2. Release to a small canary slice of live traffic with online metrics and automatic rollback guardrails.

  3. Compare canary vs control on real outcomes and cost; ramp by percentage if healthy.

  4. Keep the previous version deployable for instant rollback.


4.10 A/B significance reasoning with worked numbers

The exam does not require you to run a t-test by hand, but it does require you to know when a difference is real. The key intuition: a small sample can produce a large apparent win by chance.

Worked example — small sample. Prompt B beats A on 7 of 12 paired items (58% win rate).

text
Under the null (no difference), each item is a coin flip (p = 0.5).
Getting ≥ 7 of 12 heads by pure chance is very common (~39% two-sided).
→ 7/12 is well within noise. Do NOT promote B.

Worked example — large sample. On 2,000 paired items B wins 1,080 (54%).

text
Expected under null = 1,000; std dev ≈ sqrt(n·p·(1-p)) = sqrt(2000·0.25) ≈ 22.4
Observed excess = 1,080 − 1,000 = 80 wins → z ≈ 80 / 22.4 ≈ 3.6
z ≈ 3.6 ⇒ p < 0.001 → the 54% win is statistically significant.

Same-direction result (B better), but only the large sample supports promotion — and only if it also holds per segment (no segment regresses).

SituationCorrect action
Big win, tiny sampleCollect more data; do not promote
Small win, huge sample, significant, no segment regressionPromote via canary
Significant overall but one segment regressesDo not promote; the aggregate hides a regression
Judge is uncalibratedFix calibration before trusting any A/B result

Exam signal

Any option that says “ship B because it won N of a dozen” is the trap. The correct answer invokes larger sample + statistical significance + no per-segment regression. “MOST” and “FIRST” qualifiers usually point at “gather a significant sample” over “ship now”.


4.11 An observability and eval-record schema

To evaluate and debug at scale you need a consistent record per request. A minimal schema:

json
{
"correlation_id": "9f3a-...",
"request_id": "req_01H...",
"timestamp": "2026-09-15T11:02:33Z",
"segment": { "tenant": "42", "language": "es", "ticket_type": "refund" },
"model": "claude-sonnet-5",
"prompt_version": "support-agent@3.2.1",
"retrieval": { "k": 50, "kept": 6, "chunk_ids": ["c_101","c_233"], "recall_at_k": 0.94 },
"usage": { "input_tokens": 6120, "output_tokens": 1440, "thinking_tokens": 300 },
"cost_usd": 0.0271,
"latency_ms": { "ttft": 610, "total": 5230 },
"stop_reason": "end_turn",
"eval": { "faithfulness": 0.9, "rubric_score": 1.7, "judge_model": "claude-opus-5" },
"outcome": { "human_override": false, "user_thumb": "up" }
}
Field groupEnables
segmentPer-segment metrics (the anti-aggregate control)
prompt_version / modelAttribute regressions to a specific change; A/B attribution
retrievalDiagnose retrieval vs grounding; recall@k tracking
usage / cost_usdCost-per-task dashboards; routing decisions
latency_msp50/p95 SLA tracking (track total, watch the tail)
eval / outcomeOffline judge scores vs online human signals

If it isn’t logged, you can’t evaluate it

Segment tags, prompt/model version, retrieval detail and per-request cost are the fields most often missing — and their absence is why teams can only report a misleading aggregate. Design the schema before launch.


4.12 Scenario walkthrough: promoting a prompt change safely

Scenario. A team hand-tests a new support prompt on 15 favourite tickets, sees “clearly better” answers, and wants to ship to 100% tomorrow. The system serves English and Spanish tickets and three tenant tiers. The judge is the same Sonnet 5 session that generated the answers.

Expert reasoning trace.

  1. Reject the sample and the judge. 15 cherry-picked items prove nothing, and grading in the producing session carries its bias (anti-pattern #9). First fix the method.

  2. Build the eval set. Use a stratified golden set covering both languages and all three tiers, plus edge cases mined from production failures.

  3. Independent, calibrated judge. Grade with a separate model/session (or humans), calibrated against human labels to catch length/position bias.

  4. Paired A/B with significance. Run A vs B on the same inputs; require a statistically significant win on a large enough sample, and confirm no per-segment regression (e.g. Spanish must not drop).

  5. Gate in CI, then canary. The per-segment regression suite must pass; then canary on a small live slice with rollback before ramping.

  6. Keep the prior version hot for instant rollback.

Why the tempting alternatives are wrong: “ship because it looked better on 15” is small-sample eyeballing; “trust the same-session self-grade” is same-session bias; “flip to 100% to gather signal faster” removes the safety net; “test only English because that’s most traffic” reintroduces the aggregate-blindness that hid the Spanish failure.


4.13 Common misconceptions

MisconceptionRealityWhy it matters on the exam
“A big win on a few examples means B is better.”Small samples are dominated by noise; test for significance.Small-sample eyeballing is the classic A/B trap.
“Overall accuracy is the metric that matters.”Per-segment metrics catch failures the aggregate hides.Anti-pattern #10 appears in most D4 items.
“Mean latency reflects user experience.”Interactive SLAs live on p95/p99 tail.Mean-only distractors are wrong.
“The model can grade its own answer.”Same-session self-review carries the producing bias.Independent, calibrated judge is required.
“LLM-as-judge is objective.”Judges have length/position/self-preference bias; calibrate.Length-bias stems test this.
“Optimise cost, then measure later.”Every cost change must be re-validated per segment.Skipping re-eval is a wrong answer.
“Fix the prompt when output is wrong.”Diagnose the failing layer first (retrieval/grounding/prompt/model).Fixing the wrong layer is the expensive mistake.

Exam traps in this domain

TrapWhy it is wrong
Reporting only aggregate accuracyMasks per-segment failures (anti-pattern #10)
Using mean latency as the SLA metricHides tail; use p95
LLM-as-judge in the same session/model as the systemSame-session self-review bias (anti-pattern #9)
Trusting a judge without calibrating against humansJudge biases (position/length/self-preference) go undetected
“B looks better on 10 examples, ship it”No statistical significance; small-sample noise
Optimising cost without re-running evalsMay silently regress a segment
Fixing the prompt when retrieval was staleWrong layer; diagnose before fixing
Flipping a model change to 100% at onceNo canary; no safe rollback
Golden set skewed to one segmentAggregate looks fine while a segment fails
Batching latency-sensitive interactive traffic24 h latency violates the SLA
Promoting a variant that wins overall but regresses one segmentAggregate hides the regression; require per-segment no-regression
Trusting an A/B result from an uncalibrated judgeJudge bias contaminates the comparison; calibrate first
Omitting segment/version/cost fields from the eval recordForces misleading aggregate-only reporting
Confusing a big win rate on a tiny sample with significanceSmall samples are noise; compute/estimate significance
Tracking only total latency without TTFT for streaming UXTTFT drives perceived latency in interactive apps

Practice questions

Q1 · A support bot reports 92% overall accuracy, but complaints come only from refund-related tickets. What does the architect need? (Select one)

A. Nothing; 92% exceeds the target. B. Per-segment metrics that break accuracy down by ticket type, revealing the refund segment’s failure the aggregate hides. C. A bigger golden set of the same mix. D. A higher-effort model for all tickets.

Answer: B. Aggregate accuracy masks a segment failure (anti-pattern #10). Stratified, per-segment metrics expose the refund problem so it can be fixed. The aggregate (A) is misleading; a bigger same-mix set (C) still hides it; escalating all tickets (D) overpays without diagnosing.

Q2 · An eval uses the same model and conversation that produced the answer to grade whether the answer is correct. Why is this unsound? (Select one)

A. It costs too much. B. Same-session self-review retains the reasoning bias that produced the answer; the judge must be a separate model/session and be calibrated against humans. C. The judge should always be a bigger model. D. LLM-as-judge is never valid.

Answer: B. Grading in the producing session carries the original bias (anti-pattern #9). Independence plus human calibration is required. Cost (A) isn’t the core issue; the judge need not be bigger (C); LLM-as-judge is valid when independent and calibrated (D).

Q3 · Prompt B beats prompt A on 6 of 10 hand-picked examples. What is the correct conclusion? (Select one)

A. Ship B; it won the comparison. B. The sample is far too small to be significant; run a paired A/B on a large, per-segment eval set and test for statistical significance before promoting. C. Ship A; it is the incumbent. D. Average the two prompts.

Answer: B. A 6/10 result is noise; promotion requires a paired test on a representative sample with statistical significance and no per-segment regression. Shipping on eyeballed small samples (A), defaulting to the incumbent (C), or “averaging” prompts (D) are all unsound.

Q4 · Average latency is 2 s but users complain of slow responses. Metrics show p95 = 11 s. What should the SLA track? (Select one)

A. Mean latency only. B. Tail latency (p95/p99), since the mean hides the slow tail users actually experience. C. Total request count. D. Token count only.

Answer: B. A healthy mean with a bad p95 means the tail is the real problem; SLAs for interactive systems track p95/p99. The mean (A) hides it; request count (C) and tokens (D) are not latency SLAs.

Q5 · An LLM judge consistently rates the longer of two answers higher regardless of correctness. What is happening and the fix? (Select two)

A. Length bias in the judge. B. Calibrate against human labels and revise the rubric to score correctness/grounding explicitly, not length. C. The judge is perfectly reliable. D. Always pick the longer answer. E. Delete the eval entirely.

Answer: A and B. The judge exhibits length bias; the remedy is calibration against humans and a rubric that scores the dimensions that matter. The judge is not reliable (C), rewarding length (D) is the bug, and deleting evals (E) abandons measurement.

Q6 · A change fixed accuracy for enterprise-tier tickets but no one checked other tiers. How should this be prevented in future? (Select one)

A. Manual spot-checks after release. B. A per-segment regression suite in CI that blocks merge if any segment regresses. C. Trust the author’s judgement. D. Only test the segment that changed.

Answer: B. A per-segment regression suite in CI is exactly the guard against fixing one segment while breaking another. Manual spot-checks (A) and author trust (C) are unreliable; testing only the changed segment (D) is what caused the risk.

Q7 · A RAG answer includes a fact not present in the retrieved chunks, though recall@k is high. Which layer failed and what is the fix? (Select one)

A. Retrieval; re-index. B. Generation/grounding (hallucination); tighten the answer-only-from-context instruction, add citations, and verify faithfulness — retrieval is fine. C. Model mismatch; use Haiku. D. Infrastructure; add retries.

Answer: B. High recall means the right context was present, so an unsupported fact is a grounding/hallucination failure at generation. Re-indexing (A) targets a healthy layer; a smaller model (C) won’t help; retries (D) are unrelated.

Q8 · A team wants to cut cost 50% on a nightly bulk-summarisation job that has no latency requirement. Which lever is BEST and what must follow? (Select one)

A. Lower effort blindly and ship. B. Move the job to the Message Batches API (50% discount, within 24 h) and re-run the per-segment regression suite to confirm no quality regression. C. Switch to Haiku for all interactive traffic too. D. Disable evaluation to save compute.

Answer: B. A latency-tolerant bulk job is the Batch API’s use case; every cost change is followed by per-segment re-validation. Blind effort cuts (A) risk quality; changing interactive traffic (C) is out of scope; disabling evals (D) removes the safety net.

Q9 · A system fails only on the hardest 5% of cases and is fine elsewhere. What is the most likely cause and fix? (Select one)

A. Retrieval failure; re-chunk everything. B. Model mismatch — the tier is too small for the hardest cases; route/escalate those to a stronger model (cascade) while keeping the cheap model for the rest. C. Prompt failure; rewrite the whole prompt. D. Infrastructure; add more replicas.

Answer: B. A clean ‘fails only on the hardest cases’ signature points to model capability; a cascade escalates just those cases, preserving cost elsewhere. Re-chunking everything (A) and rewriting the prompt (C) target layers that work on the other 95%; replicas (D) don’t affect correctness.

Q10 · How should a new prompt version reach production safely? (Select two)

A. Pass the offline per-segment regression suite first. B. Canary on a small live slice with online metrics and automatic rollback, then ramp. C. Flip to 100% immediately to gather signal faster. D. Let each team edit the prompt in place. E. Skip offline evals if the change is small.

Answer: A and B. Safe rollout is offline regression gate → canary with rollback → ramp. Flipping to 100% (C) removes the safety net; in-place per-team edits (D) destroy governance; skipping offline evals for ‘small’ changes (E) is how regressions slip through.

Q11 · Which pair are the RIGHT metrics for a high-volume, cost-sensitive classification service with an interactive SLA? (Select two)

A. Cost per task. B. p95 latency. C. Mean latency only. D. Total tokens generated across the fleet. E. Number of prompt versions.

Answer: A and B. Cost-sensitive + interactive → cost per task and p95 (tail) latency are the binding metrics. Mean latency (C) hides the tail; total tokens (D) and prompt-version count (E) are not user-facing quality/SLA metrics.

Q12 · An eval dataset is 85% English support tickets; the system later fails badly on Spanish tickets in production. What was the flaw and the fix? (Select one)

A. The dataset was too small. B. The dataset lacked per-segment coverage (language); stratify the golden set so every language/segment is represented and reported separately. C. The model is broken. D. Online evals are unnecessary.

Answer: B. An unstratified set makes a whole segment invisible to offline evals. Stratifying by language (and other segments) with per-segment reporting surfaces the gap. Size (A) isn’t the issue; the model isn’t broken (C); online evals (D) are still needed but the root cause here is coverage.

Q13 · On 2,000 paired items prompt B wins 1,080 (54%); on a separate 12-item hand test B won 7. Which result should drive promotion, and why? (Select one)

A. The 12-item test, because the answers looked clearly better. B. The 2,000-item result: a 54% win at n=2,000 is ~3.6 standard deviations from chance (p < 0.001), so it is statistically significant — provided no segment regresses. C. Neither; A/B testing is unreliable. D. Average both win rates.

Answer: B. At n=2,000 the excess of 80 wins over the 1,000 expected is ≈3.6σ (std dev ≈ 22.4), which is significant; 7/12 is within coin-flip noise. The small test (A) is noise; A/B is reliable at scale (C); averaging win rates (D) is meaningless.

Q14 · A prompt change wins significantly overall but Spanish-tier accuracy drops 6 points. What is the correct decision? (Select one)

A. Promote it; the overall win is significant. B. Do not promote as-is; a per-segment regression blocks promotion even when the aggregate improves — fix the Spanish regression first. C. Promote and monitor complaints. D. Drop Spanish from the eval set.

Answer: B. Per-segment no-regression is a hard gate; an aggregate win that hides a segment regression is anti-pattern #10. Promoting anyway (A, C) ships a known regression; dropping Spanish (D) re-hides it.

Q15 · A team can only report a single overall accuracy number and cannot tell which segment fails. Which observability gap is the root cause? (Select one)

A. Missing GPU metrics. B. Eval records lack a segment tag (and prompt/model version), so metrics cannot be stratified or attributed. C. The judge model is too small. D. Latency is not logged.

Answer: B. Without segment tags and version fields, only an aggregate is computable and regressions can’t be attributed. GPU metrics (A) are irrelevant; judge size (C) doesn’t create the reporting gap; latency (D) is a different signal.

Q16 · Which TWO fields are MOST essential in a per-request eval record to support per-segment evaluation and A/B attribution? (Select two)

A. A segment tag (tenant/language/type). B. The prompt_version and model used. C. The server’s CPU temperature. D. A random UUID with no linkage. E. The marketing campaign name.

Answer: A and B. Segment tags enable stratified metrics; prompt/model version enables attributing regressions and A/B results to a specific change. CPU temperature (C), an unlinked UUID (D), and a campaign name (E) don’t support evaluation.

Q17 · Before promoting a new prompt validated only by the same session that produced the answers, what must change FIRST? (Select one)

A. Increase the model’s effort. B. Grade with an independent, calibrated judge (separate model/session) on a stratified set, because same-session self-review carries the producing bias. C. Ship to canary immediately. D. Delete the old prompt version.

Answer: B. Same-session grading is anti-pattern #9; the method must be fixed with an independent, human-calibrated judge before any promotion decision. Effort (A) doesn’t fix bias; canary (C) is premature; deleting the prior version (D) removes rollback.

Q18 · An interactive streaming UI feels slow even though total latency is acceptable. Which metric should be added? (Select one)

A. Total tokens generated. B. Time-to-first-token (TTFT), because perceived latency in streaming UIs is driven by how fast output starts. C. Daily request count. D. Number of prompt versions.

Answer: B. In streaming UIs users perceive responsiveness by when tokens start (TTFT), not just total time. Total tokens (A), request count (C), and version count (D) are not latency-experience metrics.

Key takeaways

  • Pick metrics that map to success criteria: accuracy, p95 (not mean) latency, cost per task, safety, security — and report them per segment.
  • Build stratified eval datasets (golden + synthetic + edge cases) so no segment is invisible.
  • Use rubric grading and LLM-as-judge in a separate model/session, and calibrate the judge against human labels.
  • Compare variants with paired A/B tests and require statistical significance plus no per-segment regression before promoting.
  • Run offline evals in CI as a regression gate and online evals to catch distribution shift; you need both.
  • Diagnose the failing layer (retrieval vs grounding vs prompt vs model vs segment) before fixing anything.
  • Optimise with caching, batching, routing, trimming and effort tuning — then re-validate per segment.
  • Roll out via canary with automatic rollback; never flip to 100% at once.
  • Judge significance, not vibes: a big win on 12 items is noise; a small win on thousands can be real (z ≈ excess ÷ √(n·p·(1−p))) — and it must hold per segment.
  • Design the eval record schema (segment, prompt/model version, retrieval, usage, cost, latency, judge score, outcome) before launch — you cannot evaluate what you did not log.
  • Track TTFT for streaming UX alongside p95 total latency; perceived speed is driven by first-token time.

Last updated Sep 18, 2026