AI Cert Prep
Type to search documentation.

Agents and Workflows

D5 · Reviewing and Verifying Agent Work

Verifying output you did not watch being produced — evidence trails, artefacts, spot-checking, reproducing a claim, and reviewing a long agent run efficiently.

Worth 16% — about 8 of 50 items. This domain tests the skill that makes delegation safe: verifying work you did not watch being produced. When you prompt, you see every step; when you delegate, you get a result and a story about how it was reached. This domain is about trusting the evidence, not the story — reading an evidence trail, checking artefacts, spot-checking against the definition of done, reproducing a claim, and doing all of that efficiently on a long run without re-doing the work yourself.

What you need to know

You cannot verify an unwatched agent by re-reading its confident summary — that is the story, not the evidence. Instead you verify against the definition of done (from D2), you inspect the artefacts the agent produced (files, drafts, the actual changes), and you follow the evidence trail (what it did, which sources it used, where each claim came from). You spot-check representative and high-stakes items rather than everything, and you reproduce a critical claim independently to confirm it. On a long run, you review the plan, the checkpoints and the artefacts rather than the full transcript. The governing rule: a fluent completion summary is not evidence, and “it said it did X” is not the same as “X is true”.

Learning objectives

By the end of this page you should be able to:

  1. Verify agent output against the definition of done rather than against its own summary.
  2. Distinguish the agent’s claim about what it did from evidence that it actually did it.
  3. Follow an evidence trail: sources used, steps taken, provenance of each claim.
  4. Spot-check representatively and reproduce a high-stakes claim independently.
  5. Review a long run efficiently by reading the plan, checkpoints and artefacts, not the whole transcript.
  6. Detect silent partial completion and unsupported claims.

5.1 The story is not the evidence

An agent ends a run with a summary: “I reviewed all twelve accounts, drafted responses, and everything is ready.” That sentence is a claim. It may be true, partly true, or confidently wrong. Verification means checking the claim against something independent of the claim.

The agent says…Verify by…
“I covered all twelve accounts”Counting the twelve artefacts against the list
“All figures are from the Q1 report”Opening the report and matching a sample
“I fixed every broken link”Testing a sample of links, including ones it didn’t mention
“The summary is complete”Checking it against the definition of done, item by item
text
agent's summary ──► "I did X, Y, Z" (the STORY — take as a claim)
│
▼
verification ──► artefacts + evidence trail + definition of done
(the EVIDENCE — this is what you trust)

Assessment signal

When a stem says the agent “reported success”, “summarised what it did”, or the output “looks complete”, the correct answer verifies against artefacts or the definition of done — not “trust the summary” and not “ask the agent if it’s sure” (a same-session self-check carries the same reasoning).

5.2 Verify against the definition of done

The definition of done you wrote in D2 is your verification checklist. This is why the two domains are linked: a task with no definition of done cannot be verified, because there is no acceptance test to check against. With one, review becomes mechanical — tick each criterion against the actual output.

Definition-of-done criterionVerification check
“Covers all four competitors”Count: are all four present?
“One sourced price per tier”Sample: does each price have a working source link?
“Flags anything unverifiable”Look for the flags; their absence is itself suspicious
“Fits one slide, ≤ 6 bullets”Measure it

If you find yourself unable to verify a result, the gap is usually upstream: the definition of done was too vague. Note it for the next iteration (D6).

5.3 Artefacts over assertions

An artefact is the concrete thing the agent produced or changed — the file, the draft, the actual edited document, the list of records touched. Artefacts are verifiable; assertions are not. Prefer to review the artefact directly rather than the agent’s description of it.

Assertion (weaker)Artefact (stronger)
“I updated the pricing table”The updated table itself
“I emailed the five clients”The five drafts / the sent-items entries
“I removed the outdated sections”A diff showing what was removed
“I found three issues”The three issues, each with its location

The habit to build: ask for the work product, not the report of the work product. When an agent can show the artefact, you can check it; when it can only describe it, you are back to trusting the story.

5.4 The evidence trail

A good agent run leaves a trail: the plan it followed, the tools and sources it used, and — ideally — the provenance of each claim (which source each fact came from). Following the trail is how you verify without redoing the work.

text
PLAN what it intended to do
│
STEPS/TOOLS what it actually did, in order
│
SOURCES which documents/systems it read
│
CLAIMS each output claim, linked to its source
│
ARTEFACTS the concrete outputs produced

The provenance test from foundational verification applies directly: for any claim that matters, ask where did this come from? A claim traceable to a supplied document is cheap to confirm; a claim with no source in the trail is a lead to verify, not a fact to trust.

5.5 Spot-checking and reproducing a claim

You rarely have time to verify everything, and you don’t need to. Spot-checking means checking a representative and risk-weighted sample; reproducing means independently re-deriving one high-stakes claim to confirm it.

TechniqueWhen to useHow
Representative spot-checkBulk output, uniform itemsCheck a random sample; if any fail, widen the check
Risk-weighted spot-checkMixed stakesCheck the highest-stakes items in full
Reproduce a claimA critical number or conclusionRe-derive it yourself (recompute, re-open the source)
Adversarial checkSuspicious fluencyCheck the thing it didn’t mention, or the edge case

The key spot-check discipline: a failed sample invalidates the sample. If you check five of fifty items and one is wrong, you cannot assume the other forty-five are fine — the failure means you widen the check or reject the batch, exactly as with a fabricated citation invalidating a whole output.

5.6 Reviewing a long run efficiently

A twenty-minute run can produce a transcript longer than the work would have taken you. Reading all of it defeats the purpose of delegating. Efficient review is structured, not linear.

text
Efficient long-run review (in order, stop early if a step fails):
1. Read the PLAN ── did it understand the objective?
2. Check the CHECKPOINTS ── what did it pause on / decide?
3. Verify against DONE ── tick each acceptance criterion
4. Inspect the ARTEFACTS ── the actual outputs, not the narration
5. Spot-check + reproduce── a sample + one high-stakes claim
(only read the full transcript if something above fails)

The order matters: a misread objective at step 1 means you can stop and re-brief without wading through the rest. Reading the transcript top-to-bottom is the slow, last-resort move — reserve it for when a structured check surfaces a problem you need to trace.

Assessment signal

When a stem describes a long run and asks how to review it “efficiently”, the answer is the structured path — plan, checkpoints, definition of done, artefacts, spot-check — not “read the entire transcript” and not “trust the final summary”.

Decision framework

Use the TRACE review method on any delegated result: Test against done, Reproduce a key claim, Artefacts not assertions, Check a representative sample, Evidence trail for provenance.

StepQuestionWhat you do
T — Test against doneDoes it meet every acceptance criterion?Tick the definition of done item by item
R — ReproduceIs the most critical claim actually true?Independently re-derive one high-stakes claim
A — ArtefactsAm I checking the work or the report of it?Inspect the concrete output/diff, not the summary
C — Check a sampleDo representative items hold up?Spot-check risk-weighted; a failure widens the check
E — Evidence trailWhere did each claim come from?Follow sources/provenance; unsourced claims are leads

Apply it tomorrow: on your next delegated result, run TRACE before you act on it. If any step can’t be completed, the gap is often an upstream brief defect to fix in D6.

Common mistakes

MistakeWhy it happensWhat to do instead
Trusting the completion summaryIt’s fluent and confidentVerify against artefacts and the definition of done
“Ask the agent if it’s sure”Feels like a checkSame-session self-review carries the same reasoning; check independently
Verifying nothing because it “looks complete”Polish reads as correctnessTick the acceptance criteria; look for what’s missing
Checking assertions, not artefactsThe report is easier to readInspect the actual output/diff/records
Assuming a passed sample means all passedSmall check felt sufficientA failed sample invalidates it; a passed one only covers what you checked
Reading the whole transcript to be thoroughFeels rigorousUse the structured path; transcript is last resort
Not reproducing a high-stakes numberThe agent showed workingRe-derive critical claims independently
No way to verify at allThe definition of done was vagueNote it and sharpen the brief next iteration (D6)

Scenario challenge

Scenario. Lena delegated an agent to audit her team’s 60-page knowledge base for outdated links and stale product references. Forty minutes later the agent returns a tidy report: “Reviewed all 60 pages. Found and fixed 14 broken links and updated 9 outdated product references. The knowledge base is now current.” The summary is clear and confident, and Lena is tempted to close the task. But she has an hour before a stakeholder relies on the knowledge base, and the stakes are real — customers use these pages.

Expert reasoning trace.

  1. Treat the summary as a claim. “Reviewed all 60 pages… now current” is the story. It might be true, but Lena has no evidence yet — she didn’t watch the run. Closing the task on the summary is the fluency trap.
  2. Verify against the definition of done first. Her brief asked for outdated links and stale references fixed across all pages. She checks the artefacts: does a diff or change list show edits across the pages, or only a subset? If the agent touched 40 of 60 pages, “reviewed all 60” is already suspect — a possible silent partial completion.
  3. Artefacts over assertions. Rather than trusting “fixed 14 broken links”, she looks at the actual changes — the list of links it reports fixing — and tests a sample of them. She also tests a few links it didn’t mention (the adversarial check), because a link it missed is exactly what the summary won’t surface.
  4. Spot-check, and let a failure widen the check. She samples five of the 14 “fixed” links; four resolve, one still 404s. That failed sample means she can’t trust the remaining nine by assumption — she widens to check all 14, and now distrusts the “9 references updated” claim too.
  5. Reproduce a high-stakes claim. The most consequential reference — the current pricing page link customers hit — she opens herself to confirm it points to the live page, rather than trusting the report.
  6. Follow the evidence trail for the “all 60 pages” claim. She checks the run’s step log: which pages did it actually open? If the trail shows only 44 pages read, the “all 60” claim is false and there’s an unreviewed remainder — a silent partial completion she’d never have caught from the summary.
  7. Decide and iterate. She corrects the missed links, re-scopes the unreviewed pages, and notes for D6 that the brief needed a definition of done requiring a per-page checklist and a flag for any page it couldn’t process — so the next run can’t quietly stop short.

The point. The confident summary was the least reliable part of the deliverable. Verifying against the definition of done, inspecting artefacts, spot-checking (and widening on a failure), reproducing the highest-stakes claim, and following the evidence trail caught a silent partial completion that “trust the summary” or “ask the agent if it’s sure” would have missed — and did it in far less time than re-reading 60 pages.

Assessment traps

TrapWhy it is temptingThe discriminator
“The summary says it’s done, so it’s done”Fluent summaries read as truthThe summary is a claim; verify against evidence
“Ask the agent to confirm it did the work”Feels like a verification stepSame-session self-check reuses the same reasoning; check independently
“It looks complete, no need to check”Polish signals competenceSilent partial completion looks complete; verify coverage
“Read the whole transcript to be sure”Thoroughness feels safeStructured review (plan→done→artefacts→sample) is faster and catches more
“Five samples passed, so all passed”Small check felt enoughA passed sample only covers itself; a failed one invalidates the batch
“It showed its working, so the number’s right”Working looks like proofReproduce high-stakes claims independently

Practice questions

Each item states how many responses to select. Attempt before revealing.

Q1 · An agent ends a run with 'I completed all tasks successfully.' What is the BEST next step? (Select one)

A. Close the task; the summary confirms success. B. Verify the result against the definition of done and inspect the artefacts it produced. C. Ask the agent in the same session whether it is sure. D. Re-run the whole task to compare.

Answer: B. The summary is a claim; verification means checking artefacts and the acceptance criteria independently. Trusting the summary (A) is the fluency trap. A same-session self-check (C) reuses the agent’s own reasoning. Re-running everything (D) wastes the delegation and doesn’t verify the original.

Q2 · Why is 'ask the agent if it completed the work correctly' a weak verification method? (Select one)

A. It costs too many tokens. B. A same-session self-check relies on the same reasoning that produced the result, so it can confirm its own error. C. Agents can’t answer questions about their work. D. It’s actually the strongest method.

Answer: B. Asking the agent to self-assess in the same session carries the original reasoning and bias, so it may confirm a mistake. Token cost (A) isn’t the issue. Agents can describe their work (C) — that’s exactly the unreliable claim. It is not the strongest method (D).

Q3 · What is an 'artefact' in the context of verifying agent work? (Select one)

A. The agent’s summary of what it did. B. The concrete output or change it produced — the file, draft, diff or records. C. The model’s temperature setting. D. The run’s start time.

Answer: B. An artefact is the actual work product you can inspect and check, unlike the agent’s description of it. The summary (A) is an assertion, not an artefact. Temperature (C) and start time (D) are run metadata, not the work product.

Q4 · You spot-check 5 of 50 bulk items and one fails. What should you conclude? (Select one)

A. The other 45 are fine; one failure is acceptable. B. The failed sample invalidates the assumption of uniform quality; widen the check or reject the batch. C. Re-run only the failed item and ship the rest. D. Nothing; spot-checks are unreliable.

Answer: B. A failure in the sample means you can no longer assume the unchecked items are correct, so you widen the check or reject. Assuming the rest are fine (A) ignores the signal. Fixing only the one (C) leaves 44 unverified. Spot-checks are useful, not unreliable (D).

Q5 · What is the MOST efficient way to review a long, mostly-unwatched agent run? (Select one)

A. Read the entire transcript top to bottom. B. Read the plan and checkpoints, verify against the definition of done, inspect artefacts, then spot-check — reading the full transcript only if something fails. C. Trust the final summary. D. Ask the agent to shorten its transcript.

Answer: B. Structured review — plan, checkpoints, done-criteria, artefacts, spot-check — is faster and catches more than a linear read. The full transcript (A) is the slow last resort. Trusting the summary (C) verifies nothing. Shortening the transcript (D) doesn’t verify the work.

Q6 · An agent claims 'reviewed all 60 pages' but the evidence trail shows it read only 44. What has occurred? (Select one)

A. Nothing unusual. B. A silent partial completion — it stopped short but reported full coverage; the unreviewed pages must be handled. C. A hallucinated model. D. A temperature error.

Answer: B. Reporting full coverage while the trail shows partial work is a silent partial completion, caught only by checking the evidence trail against the claim. It is not normal (A). It isn’t a ‘hallucinated model’ (C) or a temperature issue (D) — it’s a coverage-vs-claim mismatch.

Q7 · For a single high-stakes number in an agent's report, what is the strongest check? (Select one)

A. Trust it because the agent showed its working. B. Reproduce it independently — recompute or re-open the source yourself. C. Ask the agent to recompute it in the same session. D. Round it to hide any error.

Answer: B. Independently reproducing the critical claim confirms it without relying on the agent’s own reasoning. Shown working (A) can still be wrong. A same-session recompute (C) reuses the original reasoning. Rounding (D) conceals rather than checks.

Q8 · Which TWO checks best verify an agent's claim that it 'fixed all broken links'? (Select two)

A. Test a sample of the links it says it fixed. B. Test some links it did not mention, in case it missed them. C. Read its summary again more carefully. D. Ask it to confirm in the same chat. E. Assume completion because the report is detailed.

Answer: A and B. Testing a sample of the claimed fixes (A) checks the work it reported, and testing links it didn’t mention (B) is the adversarial check that catches what the summary hides. Re-reading the summary (C) and asking in-session (D) verify nothing independent. Detail (E) isn’t evidence of completion.

Q9 · Why is the definition of done central to verifying agent work? (Select one)

A. It sets the model temperature. B. It is the acceptance checklist you tick against the actual output; without it there’s nothing objective to verify against. C. It makes the agent run faster. D. It replaces the need to inspect artefacts.

Answer: B. The definition of done is the objective checklist that turns review into ticking criteria; a vague or missing one leaves nothing to verify against. It doesn’t set temperature (A) or speed (C), and it complements — not replaces — artefact inspection (D).

Q10 · An agent's report reads well but you find you cannot verify one of its claims at all. What does this most likely indicate? (Select one)

A. The claim is definitely correct. B. The brief’s definition of done was too vague to check — note it and sharpen the brief for the next iteration. C. The model is broken. D. Verification is impossible for agents.

Answer: B. An unverifiable claim usually reflects an upstream brief defect — no checkable criterion — which you fix in the next iteration (D6). Inability to verify doesn’t make the claim correct (A). It’s not a broken model (C), and agent work is verifiable when the brief supports it (D).

Q11 · Which is the difference between an assertion and an artefact you should prefer to review? (Select one)

A. Prefer the assertion — the summary is easier to read. B. Prefer the artefact — the actual updated table or diff, over the claim ‘I updated the table’. C. They are identical. D. Prefer whichever is shorter.

Answer: B. The artefact is the checkable work product; the assertion is only a claim about it, so you review the artefact. Preferring the summary (A) trusts the story. They are not identical (C), and length (D) is irrelevant to which is verifiable.

Q12 · A stakeholder will rely on an agent-produced deliverable in an hour. Which TWO actions give the most verification value in limited time? (Select two)

A. Verify against the definition of done and reproduce the single highest-stakes claim. B. Risk-weighted spot-check of the most consequential items. C. Re-read the agent’s summary twice. D. Re-run the entire task from scratch. E. Ask the agent to grade its own work.

Answer: A and B. Ticking the acceptance criteria plus reproducing the top-stakes claim (A) and risk-weighting the spot-check to the most consequential items (B) concentrate limited time where errors matter most. Re-reading the summary (C) and self-grading (E) verify nothing independent, and re-running everything (D) wastes the hour.

Key takeaways

  • The agent’s completion summary is a claim, not evidence; verify against something independent of it.
  • Verify against the definition of done — it is your acceptance checklist, and a vague one leaves nothing to check.
  • Prefer artefacts over assertions: inspect the actual output, diff or records, not the report of them.
  • Follow the evidence trail — plan, steps, sources, provenance — to verify without redoing the work.
  • Spot-check representatively and by risk; a failed sample invalidates the batch, a passed one only covers what you checked.
  • Reproduce the single highest-stakes claim independently; shown working is not proof.
  • Review a long run structurally (plan → checkpoints → done → artefacts → sample); read the full transcript only as a last resort.
  • Watch for silent partial completion — a claim of full coverage the evidence trail contradicts.

Last updated Sep 18, 2026