Agents and Workflows
D6 · Reliability and Iteration
The agent failure taxonomy, instrumenting delegations so failures are visible, and iterating a delegation brief to make it reliable rather than lucky.
Worth 16% — about 8 of 50 items. This closing domain tests whether you can make a delegation reliable, not merely successful once. A brief that worked today by luck will fail tomorrow on a slightly different input. Reliability comes from naming failures precisely (a taxonomy), making them visible (instrumentation), and iterating the brief on the real cause rather than guessing. It pulls the whole track together: every failure here maps back to a gap in the objective (D2), the context and access (D3), the boundaries (D4) or the verification (D5).
What you need to know
Agent failures fall into a small taxonomy: misunderstood objective, missing context, wrong tool or access, silent partial completion, and drift over long runs. Each has a different fix and a different upstream cause, so naming the failure is the first step to fixing it. Instrumentation means running delegations so failures are visible — capturing the plan, the evidence trail, checkpoints and outcomes — because you cannot iterate on a failure you cannot see. Iterating a delegation brief means changing one thing at a time based on the diagnosed cause: sharpen the objective, add the missing context, adjust access, add a checkpoint, or strengthen verification. The discipline is: diagnose the cause, change the brief (not the model, usually), and re-test — turning a lucky success into a repeatable one.
Learning objectives
By the end of this page you should be able to:
- Classify an agent failure using the five-mode taxonomy.
- Trace each failure mode to its upstream cause and the domain that fixes it.
- Instrument a delegation so failures are visible rather than silent.
- Iterate a brief by changing one diagnosed thing at a time and re-testing.
- Distinguish a reliable delegation from one that succeeded by luck.
- Decide when the fix is the brief, the context, the boundaries, or (rarely) the model.
6.1 The failure taxonomy
Vague debugging (“the agent didn’t work”) leads to vague fixes (“try a better model”). Precise classification leads to precise fixes. Learn these five modes cold — most exam items are asking you to name one and pick its fix.
| Failure mode | What it looks like | Upstream cause | Fixed in |
|---|---|---|---|
| Misunderstood objective | Did something adjacent to what you wanted | Goal stated as a task, or ambiguous | D2 |
| Missing context | Generic, off-brand, or wrong-fact output | Source doc / company knowledge / prior decision not supplied | D3 |
| Wrong tool or access | Stalled, or used the wrong data source | Under- or mis-provisioned tools | D3 |
| Silent partial completion | Reported “done” but did part of it | No definition of done; no coverage check | D2 / D5 |
| Drift over long runs | Started well, degraded, wandered off task | Long unwatched run with no checkpoints | D4 / D6 |
Agent failed. Which mode? ├─ wrong thing? ─► MISUNDERSTOOD OBJECTIVE → sharpen goal (D2) ├─ generic / off-brand? ─► MISSING CONTEXT → supply context (D3) ├─ stalled / wrong data? ─► WRONG TOOL/ACCESS → fix provisioning (D3) ├─ "done" but wasn't? ─► SILENT PARTIAL COMPLETION → add done-check (D2/D5) └─ degraded over time? ─► DRIFT → checkpoints/scope (D4)Assessment signal
When a stem describes a symptom, map it to the taxonomy first. “Off-brand” → missing context; “said done but wasn’t” → silent partial completion; “started fine then wandered” → drift. The right answer fixes that cause, and it’s rarely “use a bigger model”.
6.2 Silent partial completion, examined
This mode deserves its own treatment because it is the most dangerous: the agent reports success while having done only part of the work, so nothing signals a problem. It’s the failure D5’s verification exists to catch.
| Why it happens | The tell | The fix |
|---|---|---|
| No definition of done, so “done” is a feeling | Claim of coverage the evidence trail contradicts | Definition of done with an explicit coverage requirement |
| Hit a snag mid-run and skipped rather than flagged | Some items missing, no error reported | Instruct: flag what you couldn’t complete, don’t skip silently |
| Interpreted scope narrowly | Fewer outputs than the task implied | State the count/scope explicitly |
The reliability fix is a brief that makes silence impossible: “process all N items; for any you cannot complete, list it and why”. Combined with D5’s coverage check, silent partial completion becomes a caught partial completion.
6.3 Drift over long runs
On a long, unwatched run an agent can gradually wander — losing the thread of the objective, over-elaborating, or getting pulled off-task by something it found. Drift is a function of duration and lack of anchoring.
short run long run without anchoring ────────── ─────────────────────────── objective stays sharp │ objective slowly blurs output stays on-task │ output wanders / over-elaborates easy to keep on track │ each step drifts a little further │ → add checkpoints; scope the run; │ break into shorter delegationsFixes for drift are structural, not motivational: checkpoints that re-anchor to the objective, scoping the run into shorter bounded pieces, and the one→many pattern so a drifting pattern is caught after one item. Telling the agent “stay focused” is a respected boundary and won’t hold on a long run.
6.4 Instrumentation: you can’t fix what you can’t see
Instrumentation means running delegations so their behaviour is observable — the plan, the steps and tools used, the checkpoints, the evidence trail, and the outcome against the definition of done. Without it, a failure is just “it didn’t work”, and you’re guessing at the fix.
| Instrument | Makes visible | Enables |
|---|---|---|
| The plan | Whether it understood the objective | Catching misunderstood objectives early |
| Step/tool log | What it actually did and touched | Diagnosing wrong-tool and access failures |
| Checkpoints | Where it paused and what it decided | Catching drift and bad forks |
| Evidence trail | Provenance of each claim | Verification (D5) and missing-context diagnosis |
| Outcome vs done | Coverage and quality | Catching silent partial completion |
The mindset: treat a delegation like a process you can inspect, not a black box that emits an answer. When it fails, the instrumentation tells you which taxonomy mode you’re in — which tells you the fix.
Assessment signal
“How do you know why it failed?” points at instrumentation. The right answer is to examine the plan, evidence trail and outcome against the definition of done — not to re-run and hope, and not to switch models blindly.
6.5 Iterating the brief: one change at a time
Iteration is disciplined, not frantic. The rule is diagnose, change one thing, re-test. Changing several things at once means you won’t know which fixed it (or which made it worse), and reliability never accumulates.
1. OBSERVE the failure (via instrumentation) 2. CLASSIFY it (taxonomy) 3. TRACE to the upstream cause (which domain) 4. CHANGE exactly one thing in the brief/context/access/boundaries 5. RE-TEST on the same input, then on a varied one 6. repeat until reliable across varied inputs| Diagnosed cause | The one change |
|---|---|
| Misunderstood objective | Restate the goal as a checkable outcome |
| Missing context | Add the specific source / company knowledge |
| Wrong access | Grant (or scope) the specific tool |
| Silent partial completion | Add a coverage requirement to the definition of done |
| Drift | Add a checkpoint or scope the run shorter |
The model is the last lever, not the first — most failures are brief, context, access or boundary defects. Reaching for a bigger model before diagnosing is the single most common reliability anti-pattern.
6.6 Reliable versus lucky
A delegation that worked once may have worked by luck: the input happened to be easy, the ambiguity happened to resolve the way you wanted. Reliability means it works across varied inputs, including awkward ones.
| Lucky | Reliable |
|---|---|
| Worked on the one input you tried | Works across varied and edge-case inputs |
| Ambiguity resolved your way by chance | Ambiguity is specified away or gated |
| No coverage check, happened to be complete | Coverage guaranteed by the definition of done |
| Succeeded because the run was short | Anchored with checkpoints for long runs |
The test for reliability: re-run the delegation on a different, harder input. If it still meets the definition of done, the brief is reliable; if it fails, you’ve found the next thing to iterate. A single success is a hypothesis, not a result.
Decision framework
Use the LEARN loop to turn any failed or lucky delegation into a reliable one: Locate the symptom, Examine the instrumentation, Attribute to a taxonomy mode, Revise one thing, re-test on a New input.
| Step | Question | Action |
|---|---|---|
| L — Locate | What went wrong, concretely? | Describe the symptom precisely, not “it failed” |
| E — Examine | What does the trail show? | Read the plan, steps, checkpoints, outcome-vs-done |
| A — Attribute | Which failure mode is this? | Classify with the taxonomy; find the upstream cause |
| R — Revise | What single change fixes the cause? | Change one thing — brief, context, access, or boundary |
| N — New input | Is it reliable or just fixed for one case? | Re-test on a different, harder input before trusting it |
Apply it tomorrow: the next time a delegation disappoints, resist changing the model. Run LEARN, change one diagnosed thing, and prove reliability on a second input.
Common mistakes
| Mistake | Why it happens | What to do instead |
|---|---|---|
| “It didn’t work” with no classification | Debugging feels like retrying | Name the failure mode first; it points to the fix |
| Reaching for a bigger model first | Capability feels like the lever | Model is the last lever; most failures are brief/context/access |
| Changing several things at once | Wanting a fast fix | Change one thing, re-test — so you know what worked |
| Trusting a single success | It worked, so it’s done | Re-test on a varied, harder input; one success is a hypothesis |
| No instrumentation, then guessing | Black-box habit | Capture plan, trail, checkpoints, outcome vs done |
| Telling a drifting agent to “stay focused” | Instruction feels like control | Drift needs structural fixes: checkpoints, shorter scope |
| Treating silent partial completion as a fluke | It reported success | Add a coverage requirement so silence is impossible |
| Iterating without diagnosing the cause | Symptom is visible, cause isn’t | Trace to the upstream domain before changing anything |
Scenario challenge
Scenario. Sofia set up an agent on ChatGPT Work to triage her team’s incoming vendor invoices: extract the amount, match it to a purchase order, and produce a summary table for approval. It worked beautifully on her first test — one clean invoice, matched perfectly. She rolled it out for the week’s fifty invoices. The results are a mess: some amounts are wrong, several invoices are missing from the table entirely though the agent reported “all invoices processed”, and the later entries drift into inconsistent formatting and occasional editorialising about the vendors. Her instinct is to conclude agents “aren’t reliable enough for finance” and to try the most capable model.
Expert reasoning trace.
- Reject the model-first instinct. A bigger model will produce the same shaped failures faster, because none of these symptoms is a raw capability problem. Diagnose before spending.
- Classify each symptom with the taxonomy. Wrong amounts → likely missing context or wrong source (which document/field is authoritative?) — a D3 issue. Missing invoices with a “processed all” claim → silent partial completion — a D2/D5 issue. Later entries drifting into inconsistent formatting and editorialising → drift over the long run — a D4/D6 issue. Three distinct modes, three distinct fixes.
- See why the first test misled her. One clean invoice was the lucky case: no ambiguity, short run, no scale. Reliability is proven on varied inputs, and she skipped that step.
- Instrument to confirm. She examines the run: the step log shows the agent read 43 of 50 invoices (confirming silent partial completion), pulled amounts from an inconsistent field (confirming the source problem), and had no checkpoints across the long run (explaining the drift).
- Iterate one change at a time. She doesn’t rewrite everything at once. First: fix the source of truth — specify exactly which field holds the amount and which system holds the PO (D3), and re-test. Then: add a definition of done requiring all 50 processed with any it couldn’t match explicitly listed (D2/D5), and re-test. Then: add a checkpoint after the first invoice (approve the pattern) and scope the run so drift can’t accumulate (D4), and re-test.
- Prove reliability, not luck. After each change she re-tests on a varied set — a clean invoice, a messy one, one with no matching PO — until the delegation meets the definition of done across all of them. Only then is it trustworthy for the weekly run.
The point. “Agents aren’t reliable for finance” was the wrong conclusion. The delegation exhibited three separate, nameable failure modes, each with a brief-level fix; a more capable model addressed none of them. Reliability came from instrumenting the run, classifying each failure, iterating one change at a time, and proving it on varied inputs rather than trusting one lucky success.
Assessment traps
| Trap | Why it is tempting | The discriminator |
|---|---|---|
| “It’s unreliable — use the most capable model” | Capability feels like the fix | Most failures are brief/context/access/boundary defects; model is the last lever |
| “It worked in my test, so it’s reliable” | One success feels like proof | Reliability is proven on varied, harder inputs; one run is a hypothesis |
| “It said it processed everything, so it did” | Success reports reassure | Silent partial completion reports success; check coverage |
| “Tell the long-running agent to stay on task” | Instruction feels like control | Drift needs structural fixes — checkpoints, shorter scope |
| “Change several things to fix it fast” | Speed | One change at a time, or you won’t know what worked |
| “Just re-run it and hope for a better result” | Low effort | Re-running a defective brief reproduces the defect; diagnose first |
Practice questions
Each item states how many responses to select. Attempt before revealing.
Q1 · An agent produces competent but off-brand output. Which failure mode is this, and where is it fixed? (Select one)
A. Drift; add checkpoints. B. Missing context; supply company knowledge and the relevant sources (D3). C. Silent partial completion; add a coverage check. D. Wrong model; upgrade it.
Answer: B. Off-brand, generic output is the missing-context signature, fixed by supplying the company knowledge and authoritative sources in D3. It isn’t drift (A) — it didn’t degrade over time — nor partial completion (C). A model upgrade (D) won’t teach it your brand.
Q2 · An agent reports 'all 50 items processed' but 7 are missing from the output. Which failure mode is this? (Select one)
A. Drift. B. Silent partial completion. C. Missing context. D. Wrong tool.
Answer: B. Reporting full completion while some work is missing is silent partial completion, caught by a coverage check against the definition of done. Drift (A) is gradual degradation, missing context (C) is off-brand/wrong-fact output, and wrong tool (D) is a stall or wrong data source — none matches a false completeness claim.
Q3 · An agent starts a long run well but later entries wander off task and over-elaborate. What is this, and what is the correct fix? (Select one)
A. Missing context; paste more documents. B. Drift over a long run; add checkpoints and scope the run into shorter pieces. C. Wrong access; grant more tools. D. A capability limit; use a bigger model.
Answer: B. Gradual wandering on a long unwatched run is drift; the structural fix is checkpoints that re-anchor and shorter scoped runs. More documents (A), more tools (C) and a bigger model (D) don’t address a duration-driven loss of focus.
Q4 · A delegation succeeded on the one input you tried. What should you conclude? (Select one)
A. It is reliable and ready to roll out. B. It might be lucky; prove reliability by re-testing on varied, harder inputs before trusting it. C. Re-run the same input several more times. D. Switch to a cheaper model to save cost.
Answer: B. A single success is a hypothesis; reliability is demonstrated across varied and edge-case inputs. Declaring it reliable (A) risks a fragile brief. Re-running the same input (C) proves nothing new. Model/cost choices (D) don’t establish reliability.
Q5 · Why is instrumentation necessary for iterating on agent failures? (Select one)
A. It makes the agent run faster. B. You cannot fix a failure you cannot see; the plan, trail and outcome-vs-done tell you which failure mode you’re in. C. It changes the model automatically. D. It is only needed for developers writing code.
Answer: B. Instrumentation makes behaviour observable so you can classify the failure and target the fix, rather than guessing. It isn’t about speed (A) or model switching (C). It’s a delegation discipline for anyone, not only developers (D).
Q6 · When iterating a delegation brief, why change only one thing at a time? (Select one)
A. To make the process slower on purpose. B. So you can tell which change fixed the problem and reliability accumulates instead of guesswork. C. Because agents can only read one instruction. D. To use fewer tokens.
Answer: B. Changing one thing isolates cause and effect, so you learn what works and build reliability deliberately. It isn’t about slowing down (A) or token use (D), and agents read full briefs, not one instruction (C).
Q7 · In the failure taxonomy, which fix pairs correctly with 'the agent stalled and couldn't reach the CRM'? (Select one)
A. Sharpen the objective. B. Grant a scoped connector to the CRM (wrong tool/access fix). C. Add a coverage requirement. D. Use a smaller model.
Answer: B. A stall for lack of reach is a wrong-tool/access failure, fixed by provisioning the specific connector, least-privilege. Sharpening the goal (A) addresses misunderstood objectives, a coverage requirement (C) addresses partial completion, and model size (D) is unrelated to access.
Q8 · Which TWO are structural fixes for drift over a long unwatched run? (Select two)
A. Add checkpoints that re-anchor to the objective. B. Break the run into shorter, scoped delegations. C. Add ‘please stay focused’ to the brief. D. Increase the temperature. E. Remove the definition of done.
Answer: A and B. Checkpoints (A) and shorter scoped runs (B) are structural controls that stop drift accumulating. ‘Stay focused’ (C) is a respected boundary that won’t hold on a long run. Higher temperature (D) worsens variance, and removing the definition of done (E) removes your only coverage check.
Q9 · A delegation fails and a colleague immediately suggests the most capable model. What is the BEST response? (Select one)
A. Agree; capability is always the fix. B. Diagnose first — most failures are brief, context, access or boundary defects, and the model is the last lever. C. Try two models and pick the faster one. D. Abandon the task as impossible.
Answer: B. The disciplined move is to classify the failure and fix its upstream cause; the model is the last lever because it rarely addresses a vague brief or missing context. Model-first (A, C) skips diagnosis. Abandoning (D) gives up before diagnosing.
Q10 · An agent 'did something adjacent to what you wanted'. Which failure mode and fix apply? (Select one)
A. Misunderstood objective; restate the goal as a checkable outcome (D2). B. Drift; add checkpoints. C. Missing context; supply documents. D. Silent partial completion; add coverage check.
Answer: A. Producing something near-but-not the goal is a misunderstood objective, fixed by restating it as a concrete, checkable outcome. Drift (B) is gradual degradation, missing context (C) yields generic output, and partial completion (D) is a false completeness claim — none matches ‘adjacent to what you wanted’.
Q11 · A weekly invoice-triage agent shows three symptoms: wrong amounts, missing invoices reported as 'all processed', and later entries drifting. Which TWO statements reflect the correct approach? (Select two)
A. These are three distinct failure modes, each with its own brief-level fix. B. Iterate one change at a time and re-test on varied inputs after each. C. Conclude agents can’t do finance and switch to the biggest model. D. Fix all three at once by rewriting everything and shipping immediately. E. Trust the ‘all processed’ claim since it worked in the first test.
Answer: A and B. The symptoms map to three modes — wrong source, silent partial completion, and drift — each with its own fix (A), and the disciplined method is one change at a time re-tested on varied inputs (B). A bigger model (C) fixes none of the causes. Rewriting all at once (D) obscures what worked, and trusting the completion claim (E) ignores the partial completion.
Q12 · What best distinguishes a reliable delegation from a lucky one? (Select one)
A. The reliable one used a more expensive model. B. The reliable one meets the definition of done across varied, harder inputs — not just the one case that happened to be easy. C. The lucky one had a longer brief. D. There is no difference in practice.
Answer: B. Reliability is defined by holding up across varied and edge-case inputs, whereas luck is a single easy input succeeding. Model cost (A) and brief length (C) don’t determine reliability, and the distinction is very real in practice (D).
Key takeaways
- Learn the five-mode failure taxonomy — misunderstood objective, missing context, wrong tool/access, silent partial completion, drift — because naming the mode points to the fix.
- Each mode traces to an upstream domain: objective (D2), context/access (D3), boundaries (D4), verification (D5).
- Silent partial completion is the most dangerous mode; defeat it with a coverage requirement in the definition of done plus a D5 coverage check.
- Drift on long runs needs structural fixes — checkpoints and shorter scope — not “stay focused”.
- Instrument delegations so failures are visible; you cannot fix what you cannot see.
- Iterate one change at a time, tracing to the diagnosed cause; the model is the last lever, not the first.
- A single success is a hypothesis — prove reliability by re-testing on varied, harder inputs.
- Use the LEARN loop (Locate, Examine, Attribute, Revise, New input) to turn a lucky or failed delegation into a reliable one.
Last updated Sep 18, 2026