AI Cert Prep
Type to search documentation.

Agents and Workflows

Agents – Mock Exam 1

A 50-item, domain-weighted independent mock exam for the Agents and Workflows track, with full explanations and a readiness indicator.

This is a full-length, domain-weighted independent mock exam for the Agents and Workflows track. It is built from publicly available OpenAI learning objectives and is not an official OpenAI assessment. All 50 questions are new and do not repeat the domain-page items. Use this as your diagnostic: sit it first to find your two weakest domains, then revise before mock exam 2.

Instructions

  • Time: 60 minutes, matching the length of the interactive sitting.
  • Items: 50 multiple-choice and multiple-response questions. Each item states how many answers to select.
  • Selection: for a multiple-response item you must select all correct options and no incorrect ones to earn the mark; there is no partial credit.
  • No guessing penalty: answer every question — a wrong answer costs nothing beyond the mark.
  • Target: aim for at least 80% raw (40 of 50) before you take the real Academy Agents and Workflows assessment. The 80% line matches the Academy badge threshold.
  • Work each question before expanding the answer.

Domain distribution

#DomainWeightItems here
1What an Agent Is and When to Use One16%8
2Defining Objectives and Tasks18%9
3Context, Tools and Permissions18%9
4Boundaries and Guardrails16%8
5Reviewing and Verifying Agent Work16%8
6Reliability and Iteration16%8

Total: 8 + 9 + 9 + 8 + 8 + 8 = 50 items.

Readiness interpretation

This is an independent readiness indicator, not an official score.

Raw score (of 50)BandInterpretation
45–5090%+Strong readiness across all domains
40–4480–89%Assessment ready; review any weak domain
35–3970–79%Building confidence; targeted revision advised
Below 35under 70%Keep learning; revisit D2 and D3 first

The 80% band matches the Academy badge threshold, so treat 40 of 50 as your minimum before sitting the real assessment.

Take the mock exam

Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your readiness indicator, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · What an Agent Is and When to Use OneSelect one

    A finance analyst needs a single quarterly figure double-checked against one attached spreadsheet, and will read the answer before doing anything with it. Which mode fits best?

    • A. An agent with a spend cap and a web connector
    • B. A single prompt
    • C. A five-step workflow with review points
    • D. An agent that emails the figure to the team
    Show answer

    Answer: B.

    One question answered from one supplied source, read before acting, is a prompt — minimal autonomy, tools and duration. An agent (A, D) adds setup, cost and blast radius for no benefit, and the email in D is an unnecessary irreversible step. A five-step workflow (C) over-engineers a single lookup.

  2. Q2D1 · What an Agent Is and When to Use OneSelect one

    Which characteristic most clearly signals that a task genuinely needs an agent rather than a workflow?

    • A. The task is emotionally important to the requester
    • B. The next step depends on what earlier steps discover, so the path cannot be scripted in advance
    • C. The task involves more than 200 words of instructions
    • D. The requester is a senior manager
    Show answer

    Answer: B.

    An agent earns its keep when the path cannot be fixed ahead of time, so it must plan as it learns. Emotional weight (A), instruction length (C) and requester seniority (D) do not change whether the steps are knowable in advance, which is what separates a workflow from an agent.

  3. Q3D1 · What an Agent Is and When to Use OneSelect one

    An operations lead schedules an identical month-end reconciliation with the same fixed steps every month and reviews between each step. Which mode is most appropriate?

    • A. A prompt
    • B. A workflow
    • C. An agent that plans its own steps
    • D. An agent with broad connector access
    Show answer

    Answer: B.

    Fixed, known steps with review points between them are the definition of a workflow, which is cheaper to build and easier to verify. A single prompt (A) cannot carry a sequenced multi-step process, and an agent (C, D) adds autonomy that a fully known path does not require.

  4. Q4D1 · What an Agent Is and When to Use OneSelect one

    Why does a long, unwatched agent run demand different oversight than a short prompt you watch?

    • A. Long runs always cost more money
    • B. You cannot stop a wrong turn live, so trust must come from designed boundaries and after-the-fact verification
    • C. Long runs always use a larger model
    • D. Short prompts cannot use tools
    Show answer

    Answer: B.

    Watched work is self-correcting because you can halt a bad step; an unwatched run has already acted by the time you look, so control shifts to boundaries and verification. Cost (A) and model size (C) are not the defining difference, and short prompts can use tools you invoke (D).

  5. Q5D1 · What an Agent Is and When to Use OneSelect one

    A colleague says 'this report is really complicated, so it obviously needs an agent'. What is the soundest reply?

    • A. Agree — complicated work is exactly what agents are for
    • B. Difficulty alone does not decide the mode; judge path, access, duration and reversibility instead
    • C. Split it into fifty prompts regardless
    • D. Use the most capable model available for anything complex
    Show answer

    Answer: B.

    Fit is decided by the four axes, not by how hard a task feels; a complicated task with a fixed path is still a workflow. Treating difficulty as the deciding factor (A) is the trap, blindly splitting (C) ignores whether steps are fixed, and model choice (D) is unrelated to the mode.

  6. Q6D1 · What an Agent Is and When to Use OneSelect one

    A knowledge worker wants to hand off compiling a weekly market-news digest whose sources vary week to week, running unattended for about ten minutes and producing only a draft. Which mode fits and why?

    • A. A prompt, because it only produces a draft
    • B. A workflow, because the steps never change
    • C. An agent, because the path cannot be scripted and it runs unattended toward a reversible draft
    • D. A single very long prompt covering every possible source
    Show answer

    Answer: C.

    Varying sources mean the path is discovered as it goes, and an unattended multi-step run producing a reversible draft is agent territory. It is not a prompt (A, D) because it is not one watched step, and not a workflow (B) because the steps are not fixed.

  7. Q7D1 · What an Agent Is and When to Use OneSelect two

    Which TWO situations most clearly make an agent the wrong tool for the job?

    • A. You cannot state a clear definition of done for the task
    • B. The task benefits from several tools used over an unwatched run
    • C. Verifying the unwatched output would take longer than doing the task yourself
    • D. The path depends on what earlier steps discover
    • E. The actions are reversible drafts inside the workspace
    Show answer

    Answer: A and C.

    A missing definition of done (A) means you cannot brief or verify the work, and verification costing more than the work (C) makes delegation a net loss — both are classic 'wrong tool' signals. Multi-tool unwatched runs (B) and discovery-driven paths (D) point toward an agent, and reversible drafts (E) simply keep the risk low.

  8. Q8D1 · What an Agent Is and When to Use OneSelect one

    What best captures what 'autonomy' means when distinguishing an agent from a detailed workflow?

    • A. How many words the instructions contain
    • B. Who chooses the next step — you, or the system
    • C. How fast the response streams back
    • D. Whether the output is a table or prose
    Show answer

    Answer: B.

    Autonomy is about who decides the path: in a workflow you sequence the steps, whereas an agent plans and chooses them. Instruction length (A) measures size not autonomy, and streaming speed (C) and output format (D) are unrelated.

  9. Q9D2 · Defining Objectives and TasksSelect one

    Which delegated objective is most ready to hand to an agent?

    • A. Help sort out the onboarding situation
    • B. Have a look at the survey responses and see what stands out
    • C. Produce a one-page summary of the three most common complaints in the attached survey export, with a count and one verbatim example each
    • D. Make onboarding better somehow
    Show answer

    Answer: C.

    Option C states a concrete outcome, a bounded scope and a named source of truth, so the agent can execute and you can verify. A, B and D are wishes with no defined artefact or acceptance test.

  10. Q10D2 · Defining Objectives and TasksSelect two

    Which TWO properties make a definition of done useful for both the agent and your later verification?

    • A. It is observable, so you can tick each criterion without a judgment call
    • B. It is failure-aware, telling the agent to flag what it cannot meet rather than fabricate it
    • C. It names the model and reasoning effort to use
    • D. It sets the maximum run time in minutes
    • E. It lists every connector the agent may open
    Show answer

    Answer: A and B.

    An observable definition of done (A) lets you verify by ticking criteria, and a failure-aware one (B) surfaces gaps instead of hiding fabrication — both serve execution and verification. Model settings (C), a run time (D) and a connector list (E) are other parts of a brief, none of which is what makes 'done' checkable.

  11. Q11D2 · Defining Objectives and TasksSelect one

    A brief says 'never quote a supplier's confidential pricing' and 'try to keep the memo under one page'. How should the agent treat these two instructions?

    • A. Both are suggestions it may drop under time pressure
    • B. The confidentiality rule is a hard constraint it must not violate; the length target is a preference
    • C. Both are hard constraints of equal weight
    • D. Both are preferences
    Show answer

    Answer: B.

    A confidentiality rule stated as 'never' is a hard constraint, while a 'try to keep' length target is a preference it can exceed when needed. Treating the confidentiality rule as optional (A, D) is unsafe, and treating the length target as hard (C) could block otherwise good work.

  12. Q12D2 · Defining Objectives and TasksSelect one

    An agent blends this year's headcount with a figure that turns out to be two years old. What is the most likely root cause and fix?

    • A. Reasoning effort too high; lower it
    • B. No source of truth named; specify and rank the authoritative sources
    • C. The model is too small; upgrade it
    • D. The brief was too short; make it longer for its own sake
    Show answer

    Answer: B.

    Blending stale and current data is the classic missing-source-of-truth symptom, fixed by naming which source is authoritative and how to resolve conflicts. Reasoning effort (A) governs how hard it thinks, model size (C) and mere length (D) do not tell it which figure to trust.

  13. Q13D2 · Defining Objectives and TasksSelect one

    A brief asks an agent to apply a new naming convention across 80 files. Where does the single highest-value checkpoint usually belong?

    • A. Only after all 80 files are renamed
    • B. After the first file, so you can approve the pattern before the other 79 run
    • C. Every 15 seconds regardless of progress
    • D. Never — checkpoints only slow the agent down
    Show answer

    Answer: B.

    The 'one before many' checkpoint converts an 80-way mistake into a one-way one by approving the approach before it scales. Reviewing only at the end (A) means the error already repeated, time-based pauses (C) do not align with risk, and no checkpoints (D) remove your steering.

  14. Q14D2 · Defining Objectives and TasksSelect one

    An objective is ambiguous on one specific, identifiable point — which of two teams the report is 'for'. What is the best instruction to the agent?

    • A. Pick one interpretation silently and keep moving
    • B. Pause and ask which team before proceeding
    • C. Skip the ambiguous part entirely
    • D. Write two full reports, one for each team
    Show answer

    Answer: B.

    A single identifiable fork is best handled by a checkpoint: pause and ask before the assumption propagates. Deciding silently (A) risks a wrong path, skipping it (C) leaves the task incomplete, and producing two full versions (D) is wasteful when one question resolves it.

  15. Q15D2 · Defining Objectives and TasksSelect one

    An agent could not find a figure it needed, so it produced a plausible-looking number to complete the table. What should the brief have instructed?

    • A. Always fill in missing values so the output looks complete
    • B. Flag any value it cannot verify rather than fabricating one
    • C. Abandon the whole task at the first missing value
    • D. Use a larger model next time
    Show answer

    Answer: B.

    A failure-aware definition of done tells the agent to surface gaps instead of inventing data. Filling gaps (A) hides fabrication, halting the entire task (C) is disproportionate when one gap can be flagged, and model size (D) does not change fabrication behaviour.

  16. Q16D6 · Reliability and IterationSelect one

    In the five-mode failure taxonomy, which mode is described as the most dangerous because the agent reports success while having done only part of the work?

    • A. Misunderstood objective
    • B. Silent partial completion
    • C. Missing context
    • D. Drift over long runs
    Show answer

    Answer: B.

    Silent partial completion is the most dangerous mode because nothing signals a problem — the agent claims full coverage while some work is missing. Misunderstood objective (A) produces adjacent work, missing context (C) produces off-brand or wrong-fact output, and drift (D) is gradual degradation over a long run — each is visible in a different way.

  17. Q17D2 · Defining Objectives and TasksSelect two

    Which TWO items belong in a strong definition of done for a slide-ready competitor summary?

    • A. Covers all four named competitors
    • B. Uses the newest available model
    • C. Fits one slide with no more than six bullets
    • D. Completes in under five minutes
    • E. Uses an upbeat, marketing tone throughout
    Show answer

    Answer: A and C.

    Coverage of all four competitors (A) and a size bound of one slide with six bullets (C) are observable acceptance criteria you can tick. Model choice (B) and run time (D) are operational, and tone (E) is a preference rather than an acceptance test for a data summary.

  18. Q18D2 · Defining Objectives and TasksSelect one

    Why is 'state the outcome, not the keystrokes' the recommended way to write a goal?

    • A. Steps are always shorter than outcomes
    • B. Describing your own steps constrains the agent to your path and hides the acceptance test, while an outcome lets it plan and stays checkable
    • C. Outcomes always require a bigger model
    • D. Agents cannot follow step-by-step instructions at all
    Show answer

    Answer: B.

    An outcome lets the agent plan the path while remaining verifiable, whereas a list of your keystrokes constrains it and obscures what 'done' means. Length (A) is not the point, outcomes do not require larger models (C), and agents can follow steps (D) — they simply plan better from an outcome.

  19. Q19D3 · Context, Tools and PermissionsSelect one

    An agent returns competent but generic, off-brand copy. Which fix is most appropriate?

    • A. Grant it more connectors and tools
    • B. Supply the brand voice, style guide and product facts as authoritative company knowledge
    • C. Switch to a larger model
    • D. Give it write access to the website
    Show answer

    Answer: B.

    Off-brand output is a missing-context symptom: the agent lacks the company knowledge that defines the brand. More tools (A) and write access (D) add access rather than context, and a larger model (C) still will not know your brand without being told.

  20. Q20D3 · Context, Tools and PermissionsSelect one

    A task only requires the agent to answer questions from one uploaded policy document. What access should it receive?

    • A. Web search plus write access to the shared drive
    • B. Read access to that document and nothing more
    • C. A connector to every company system, to be safe
    • D. Send-email capability so it can share answers
    Show answer

    Answer: B.

    Least privilege means granting exactly the read access the task needs. Web and write (A), broad connectors (C) and send (D) all add blast radius the task never requires.

  21. Q21D3 · Context, Tools and PermissionsSelect one

    Which set best names the four kinds of context an agent typically needs?

    • A. Task brief, source documents, company knowledge, prior decisions
    • B. Temperature, top-p, max tokens, seed
    • C. Model, region, plan tier, language
    • D. Font, colour, layout, length
    Show answer

    Answer: A.

    The four context types are the task brief, the source documents for the task, relevant company knowledge and prior decisions already settled. B lists sampling settings, C lists account facts and D lists formatting — none is the context an agent reasons from.

  22. Q22D3 · Context, Tools and PermissionsSelect one

    You connect an agent to an entire shared drive so it can open one specific file. What is the risk?

    • A. None; it will only open the file you meant
    • B. The agent can reach every file in the drive and may pull from files you never intended
    • C. It will run more slowly
    • D. It will silently switch models
    Show answer

    Answer: B.

    A connector exposes everything it reaches, and an autonomous agent may use any of it toward the goal, surfacing data far beyond your intended file. It will not politely restrict itself (A), and speed (C) and model choice (D) are unaffected by scope.

  23. Q23D3 · Context, Tools and PermissionsSelect one

    An agent reports 'I cannot access the CRM' and stops. What is the correct fix?

    • A. Paste more background documents into the task
    • B. Grant a scoped read connector to the CRM records the task needs
    • C. Use a bigger model
    • D. Lower the reasoning effort
    Show answer

    Answer: B.

    'Cannot access' is a missing-access problem, fixed by granting the specific connector scoped to the records required. Pasting documents (A) addresses context, and model size (C) or reasoning effort (D) do not grant access.

  24. Q24D3 · Context, Tools and PermissionsSelect one

    Why can supplying too much context actually hurt an agent's output?

    • A. It never hurts; more context is always better
    • B. Excess and stale documents bury the relevant signal and can anchor the agent on the wrong source
    • C. It changes the model's price tier
    • D. It disables the agent's tools
    Show answer

    Answer: B.

    Dumping everything in buries the authoritative signal and can make the agent over-weight irrelevant or outdated material. More is not always better (A), and it neither changes pricing (C) nor disables tools (D).

  25. Q25D3 · Context, Tools and PermissionsSelect one

    A task must produce an email draft you will send yourself after review. How should access be arranged for least privilege?

    • A. Give it autonomous send so it can finish end to end
    • B. Give it draft creation now, with the send tool available only behind an approval checkpoint
    • C. Give it write access to the whole mail system
    • D. Give it no tools; it can just describe the email in prose
    Show answer

    Answer: B.

    Draft-before-send with the irreversible send gated behind approval is the least-privilege, safe arrangement. Autonomous send (A) removes the gate, whole-system write (C) is far broader than needed, and no tools (D) under-provisions a task that must produce a draft.

  26. Q26D3 · Context, Tools and PermissionsSelect one

    An agent keeps reopening a question the team settled last quarter. Which kind of context was missing?

    • A. Prior decisions — what has already been settled
    • B. A larger context window
    • C. Web search access
    • D. Write access to the project
    Show answer

    Answer: A.

    Relitigating a settled question is the signature of missing 'prior decisions' context; supply the record of what was decided. A bigger window (B) will not help if the decision was never provided, and web (C) and write (D) are access, not the missing context.

  27. Q27D3 · Context, Tools and PermissionsSelect two

    Which TWO moves best reduce an agent's blast radius without reducing its ability to do a read-only research task?

    • A. Grant read access instead of write
    • B. Scope a connector to the specific folder rather than the whole drive
    • C. Enable every connector available
    • D. Turn on autonomous send for convenience
    • E. Remove the definition of done
    Show answer

    Answer: A and B.

    Read-instead-of-write (A) and scoping a connector to the needed folder (B) shrink blast radius while leaving a read-only task fully doable. Every connector (C) and autonomous send (D) enlarge blast radius, and removing the definition of done (E) harms the task without improving safety.

  28. Q28D6 · Reliability and IterationSelect one

    An agent worked perfectly on the one clean input tested but failed on the real batch. Which upstream lesson does this teach about proving a delegation?

    • A. One clean success is proof the brief is reliable
    • B. A single easy success can be luck; reliability must be proven on varied, harder inputs before rollout
    • C. The model was simply too small on the batch
    • D. The batch was too large to ever delegate
    Show answer

    Answer: B.

    The clean input was the lucky case; reliability is only established by re-testing on varied and edge-case inputs. Treating one success as proof (A) is exactly the trap, model size (C) rarely explains a brief that was never stress-tested, and batch size alone (D) does not make delegation impossible.

  29. Q29D4 · Boundaries and GuardrailsSelect one

    Which action most requires an enforced approval gate before it can fire?

    • A. Drafting a summary in the workspace
    • B. Posting a message in an external customer channel
    • C. Analysing a spreadsheet you supplied
    • D. Rewriting a paragraph you will read next
    Show answer

    Answer: B.

    Posting to an external customer channel is irreversible, so it needs an enforced gate before it fires. Drafting (A), analysing (C) and rewriting (D) are reversible and safe with light review because nothing has left the workspace.

  30. Q30D4 · Boundaries and GuardrailsSelect one

    What is the key difference between a respected boundary and an enforced boundary?

    • A. Respected boundaries are simply written in a firmer tone
    • B. A respected boundary asks the agent not to cross a line; an enforced boundary makes crossing it impossible
    • C. Enforced boundaries only apply to large models
    • D. There is no meaningful difference
    Show answer

    Answer: B.

    A respected boundary is instruction the agent usually honours, whereas an enforced boundary is a system control that removes the possibility of crossing. Tone (A) is irrelevant, enforcement is not model-specific (C), and the difference is central (D).

  31. Q31D4 · Boundaries and GuardrailsSelect one

    For an irreversible action, why is 'just instruct the agent clearly not to do it' insufficient?

    • A. It is sufficient; clear instructions always hold
    • B. A misread objective or unexpected path can lead the agent across a merely-requested line, and the action cannot be undone
    • C. Instructions cost too many tokens
    • D. Agents ignore every instruction they are given
    Show answer

    Answer: B.

    Instruction depends on correct understanding, and an autonomous agent can cross a requested line by misreading the goal — unacceptable when the action is irreversible. Clear instructions do not always hold (A), token cost (C) is irrelevant, and agents do not ignore all instructions (D) — they can misinterpret them.

  32. Q32D4 · Boundaries and GuardrailsSelect one

    An agent that can iterate and call paid tools is told 'please don't spend too much'. What is the problem?

    • A. Nothing; the instruction is enough
    • B. That is a respected boundary that will not stop a looping run; set an enforced spend cap
    • C. Paid tools cannot loop
    • D. Spend limits slow the model down
    Show answer

    Answer: B.

    A vague requested limit will not halt a runaway; an enforced spend cap makes the worst case bounded and known. The instruction alone is not enough (A), loops can occur with paid tools (C), and an enforced cap stops a runaway rather than slowing the model (D).

  33. Q33D4 · Boundaries and GuardrailsSelect one

    What is the STRONGEST data boundary against an agent leaking a sensitive internal document?

    • A. Instructing it not to share the document
    • B. Not granting the agent reach to the sensitive document at all
    • C. Asking it to summarise the document carefully
    • D. Using a larger model
    Show answer

    Answer: B.

    The strongest data boundary is architectural: an agent cannot leak what it was never able to reach. Instruction (A) can be missed, careful summarisation (C) still touches the data, and model size (D) creates no data boundary.

  34. Q34D4 · Boundaries and GuardrailsSelect one

    Why can gating every single action be as harmful as gating none?

    • A. It is not; more gates are always safer
    • B. Excessive gates lead humans to approve reflexively, so a real risk slips through unnoticed
    • C. Gates disable the agent's tools
    • D. Gates silently change the model
    Show answer

    Answer: B.

    Over-gating produces rubber-stamp fatigue, defeating the gate when it matters. More gates are not always better (A), and gates neither disable tools (C) nor change the model (D).

  35. Q35D4 · Boundaries and GuardrailsSelect two

    Which TWO boundaries should be system-enforced rather than merely written in the brief?

    • A. A preference for bullet points over prose
    • B. An approval gate before any external send
    • C. A hard spend cap on a task that uses paid tools
    • D. A suggestion to keep the tone friendly
    • E. A wish to finish before the end of the day
    Show answer

    Answer: B and C.

    An approval gate on irreversible sends (B) and a hard spend cap on paid-tool use (C) protect against real, hard-to-reverse harm and must be enforced. Style (A), tone (D) and a soft timing wish (E) are preferences for which respected boundaries are fine.

  36. Q36D4 · Boundaries and GuardrailsSelect one

    An agent 'improving' a document overwrote the original, losing the prior version. Which boundary would have prevented this?

    • A. A friendlier tone in the brief
    • B. An enforced gate on overwriting originals, or forcing a save to a new version
    • C. A larger context window
    • D. A web-search tool
    Show answer

    Answer: B.

    Overwriting can be irreversible, so gating overwrites or forcing versioned saves preserves the original. Tone (A), context window (C) and web search (D) do nothing to protect the file.

  37. Q37D5 · Reviewing and Verifying Agent WorkSelect one

    An agent ends a run with 'I completed every task successfully.' What is the best next step?

    • A. Close the task; the summary confirms success
    • B. Verify the result against the definition of done and inspect the artefacts it produced
    • C. Ask the agent in the same session whether it is sure
    • D. Re-run the whole task to compare
    Show answer

    Answer: B.

    The summary is a claim; verification means checking artefacts and acceptance criteria independently. Trusting the summary (A) is the fluency trap, a same-session self-check (C) reuses the agent's own reasoning, and re-running everything (D) wastes the delegation without verifying the original.

  38. Q38D5 · Reviewing and Verifying Agent WorkSelect one

    Why is 'ask the agent whether it did the work correctly' a weak verification method?

    • A. It uses too many tokens
    • B. A same-session self-check relies on the same reasoning that produced the result, so it can confirm its own error
    • C. Agents cannot describe their own work
    • D. It is actually the strongest method
    Show answer

    Answer: B.

    Asking the agent to self-assess in the same session carries the original reasoning and bias, so it may confirm a mistake. Token cost (A) is not the issue, agents can describe their work (C) — that is the unreliable claim — and it is not the strongest method (D).

  39. Q39D5 · Reviewing and Verifying Agent WorkSelect one

    In verifying agent work, what is an 'artefact'?

    • A. The agent's summary of what it did
    • B. The concrete output or change it produced — the file, draft, change list or records
    • C. The reasoning-effort setting
    • D. The run's start time
    Show answer

    Answer: B.

    An artefact is the actual work product you can inspect and check, unlike the agent's description of it. The summary (A) is an assertion, and effort setting (C) and start time (D) are run metadata, not the work product.

  40. Q40D5 · Reviewing and Verifying Agent WorkSelect one

    You spot-check 6 of 60 bulk items and one fails. What should you conclude?

    • A. The other 54 are fine; one failure is acceptable
    • B. The failed sample invalidates the assumption of uniform quality; widen the check or reject the batch
    • C. Fix only the failed item and ship the rest
    • D. Nothing; spot-checks are unreliable
    Show answer

    Answer: B.

    A failure in the sample means you can no longer assume the unchecked items are correct, so you widen the check or reject. Assuming the rest are fine (A) ignores the signal, fixing only one (C) leaves 53 unverified, and spot-checks are useful, not unreliable (D).

  41. Q41D5 · Reviewing and Verifying Agent WorkSelect one

    What is the most efficient way to review a long, mostly-unwatched agent run?

    • A. Read the entire transcript top to bottom
    • B. Read the plan and checkpoints, verify against the definition of done, inspect artefacts, then spot-check — reading the full transcript only if something fails
    • C. Trust the final summary
    • D. Ask the agent to shorten its transcript
    Show answer

    Answer: B.

    Structured review — plan, checkpoints, done-criteria, artefacts, spot-check — is faster and catches more than a linear read. The full transcript (A) is a last resort, trusting the summary (C) verifies nothing, and shortening the transcript (D) does not verify the work.

  42. Q42D5 · Reviewing and Verifying Agent WorkSelect one

    For a single high-stakes number in an agent's report, what is the strongest check?

    • A. Trust it because the agent showed its working
    • B. Reproduce it independently — recompute it or re-open the source yourself
    • C. Ask the agent to recompute it in the same session
    • D. Round it to hide any small error
    Show answer

    Answer: B.

    Independently reproducing the critical claim confirms it without relying on the agent's own reasoning. Shown working (A) can still be wrong, a same-session recompute (C) reuses the original reasoning, and rounding (D) conceals rather than checks.

  43. Q43D5 · Reviewing and Verifying Agent WorkSelect two

    Which TWO checks best verify an agent's claim that it 'updated every out-of-date figure' in a document?

    • A. Sample some of the figures it says it updated and confirm them against the source
    • B. Check a few figures it did not mention, in case it missed them
    • C. Read its summary again more carefully
    • D. Ask it to confirm in the same chat
    • E. Assume completion because the report is detailed
    Show answer

    Answer: A and B.

    Sampling the claimed updates (A) checks the reported work, and checking figures it did not mention (B) is the adversarial check that catches what the summary hides. Re-reading the summary (C) and asking in-session (D) verify nothing independent, and detail (E) is not evidence of completion.

  44. Q44D5 · Reviewing and Verifying Agent WorkSelect one

    You find you cannot verify one of an agent's claims at all against any evidence. What does this most likely indicate?

    • A. The claim is definitely correct
    • B. The brief's definition of done was too vague to check — note it and sharpen the brief next iteration
    • C. The model is broken
    • D. Verification is simply impossible for agents
    Show answer

    Answer: B.

    An unverifiable claim usually reflects an upstream brief defect — no checkable criterion — which you fix in the next iteration. Inability to verify does not make the claim correct (A), it is not a broken model (C), and agent work is verifiable when the brief supports it (D).

  45. Q45D6 · Reliability and IterationSelect two

    An agent reports 'all 40 records processed' but 6 are missing from the output. Which TWO responses are correct?

    • A. Classify it as silent partial completion — a claim of full coverage the output contradicts
    • B. Add a coverage requirement to the definition of done so the agent must list anything it could not complete
    • C. Trust the completion claim because the report was detailed
    • D. Switch to a larger model to fix the missing records
    • E. Remove the definition of done to simplify the next run
    Show answer

    Answer: A and B.

    Reporting full completion while work is missing is silent partial completion (A), and the reliability fix is a coverage requirement that forces the agent to flag what it could not do (B). Trusting the claim (C) ignores the mismatch, a bigger model (D) does not address a missing coverage check, and removing the definition of done (E) removes your only means of catching it.

  46. Q46D6 · Reliability and IterationSelect one

    An agent starts a long run well, then later entries wander off task and over-elaborate. What is this and the correct fix?

    • A. Missing context; paste more documents
    • B. Drift over a long run; add checkpoints and scope the run into shorter pieces
    • C. Wrong access; grant more tools
    • D. A capability limit; use a bigger model
    Show answer

    Answer: B.

    Gradual wandering on a long unwatched run is drift; the structural fix is checkpoints that re-anchor plus shorter scoped runs. More documents (A), more tools (C) and a bigger model (D) do not address a duration-driven loss of focus.

  47. Q47D6 · Reliability and IterationSelect one

    A delegation succeeded on the one input you tried. What should you conclude?

    • A. It is reliable and ready to roll out
    • B. It might be lucky; prove reliability by re-testing on varied, harder inputs before trusting it
    • C. Re-run the same input several more times
    • D. Switch to a cheaper model to save cost
    Show answer

    Answer: B.

    A single success is a hypothesis; reliability is demonstrated across varied and edge-case inputs. Declaring it reliable (A) risks a fragile brief, re-running the same input (C) proves nothing new, and cost choices (D) do not establish reliability.

  48. Q48D6 · Reliability and IterationSelect one

    Why change only one thing at a time when iterating a delegation brief?

    • A. To slow the process down deliberately
    • B. So you can tell which change fixed the problem and reliability accumulates instead of guesswork
    • C. Because agents can only read one instruction
    • D. To use fewer tokens
    Show answer

    Answer: B.

    Changing one thing isolates cause and effect, so you learn what works and build reliability deliberately. It is not about slowing down (A) or token use (D), and agents read full briefs, not one instruction (C).

  49. Q49D6 · Reliability and IterationSelect one

    A delegation fails and a colleague immediately suggests switching to the most capable model. What is the best response?

    • A. Agree; capability is always the fix
    • B. Diagnose first — most failures are brief, context, access or boundary defects, and the model is the last lever
    • C. Try two models and pick the faster one
    • D. Abandon the task as impossible
    Show answer

    Answer: B.

    The disciplined move is to classify the failure and fix its upstream cause; the model is the last lever because it rarely addresses a vague brief or missing context. Model-first thinking (A, C) skips diagnosis, and abandoning (D) gives up before diagnosing.

  50. Q50D6 · Reliability and IterationSelect two

    Which TWO are structural fixes for drift over a long unwatched run?

    • A. Add checkpoints that re-anchor to the objective
    • B. Break the run into shorter, scoped delegations
    • C. Add 'please stay focused' to the brief
    • D. Increase the reasoning effort
    • E. Remove the definition of done
    Show answer

    Answer: A and B.

    Checkpoints (A) and shorter scoped runs (B) are structural controls that stop drift accumulating. 'Stay focused' (C) is a respected boundary that will not hold on a long run, higher effort (D) does not address duration-driven wandering, and removing the definition of done (E) removes your coverage check.

Last updated Sep 18, 2026