Agents and Workflows
Agents – Mock Exam 2
A harder 50-item, domain-weighted independent mock exam for the Agents and Workflows track, designed as a timed readiness gate.
This is the second full-length independent mock exam for the Agents and Workflows track, built from publicly available OpenAI learning objectives and not an official OpenAI assessment. It is deliberately harder than mock exam 1 — more multi-constraint stems and more FIRST / BEST / MOST cost-effective / TWO qualifiers. All 50 questions are new and distinct from both mock exam 1 and the domain-page items. Use this as your readiness gate under timed conditions once mock exam 1 and your revision are behind you.
Instructions
- Time: 60 minutes. Sit this one timed, in a single uninterrupted block.
- Items: 50 multiple-choice and multiple-response questions. Each item states how many answers to select.
- Selection: for a multiple-response item you must select all correct options and no incorrect ones to earn the mark; there is no partial credit.
- No guessing penalty: answer every question — a wrong answer costs nothing beyond the mark.
- Target: aim for at least 80% raw (40 of 50) under timed conditions before you take the real Academy Agents and Workflows assessment. The 80% line matches the Academy badge threshold.
- Read multi-constraint stems carefully: the qualifier (
FIRST,BEST,TWO) usually decides between two defensible options.
Domain distribution
| # | Domain | Weight | Items here |
|---|---|---|---|
| 1 | What an Agent Is and When to Use One | 16% | 8 |
| 2 | Defining Objectives and Tasks | 18% | 9 |
| 3 | Context, Tools and Permissions | 18% | 9 |
| 4 | Boundaries and Guardrails | 16% | 8 |
| 5 | Reviewing and Verifying Agent Work | 16% | 8 |
| 6 | Reliability and Iteration | 16% | 8 |
Total: 8 + 9 + 9 + 8 + 8 + 8 = 50 items.
Readiness interpretation
This is an independent readiness indicator, not an official score.
| Raw score (of 50) | Band | Interpretation |
|---|---|---|
| 45–50 | 90%+ | Strong readiness; you are ready for the real assessment |
| 40–44 | 80–89% | Assessment ready; review any weak domain |
| 35–39 | 70–79% | Building confidence; another revision pass advised |
| Below 35 | under 70% | Keep learning; this harder set has found real gaps |
Because this mock is harder than mock exam 1, a score in the 80–89% band here is a strong signal of readiness. Treat 40 of 50 as your minimum gate.
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your readiness indicator, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A team lead has a task whose steps are fully known and identical each run, but which touches an irreversible external send. Considering path, access, duration and reversibility together, which build is MOST appropriate?
Show answer
Answer: B.
A fixed, known path points to a workflow rather than an agent, and the irreversible send must sit behind a human gate. Autonomous send (A) removes the only control that matters, a single prompt (C) cannot carry a sequenced process, and a broad-access agent (D) adds autonomy the fixed path does not need plus unnecessary blast radius.
Two tasks look similar in difficulty. Task X has steps that depend on what earlier steps find and runs unattended; task Y has a fixed sequence you review between. Which is the BEST classification?
Show answer
Answer: B.
Fit is decided by whether the path is knowable in advance: X must plan (agent), Y is fixed with reviews (workflow). Treating difficulty as decisive (A) is the trap, both being prompts (C) ignores their multi-step nature, and D reverses the correct mapping.
A manager wants to automate a recurring task end to end, including a public publish step, arguing it is the most tedious part. What is the FIRST thing you should establish?
Show answer
Answer: B.
Reversibility governs whether automation is safe: an un-gateable irreversible step is a signal not to automate that step. Model cost (A) and brief length (C) are secondary, and what a competitor does (D) is irrelevant to whether this action can be safely delegated.
A stakeholder insists that because a task will take an agent twenty unattended minutes, it must be 'more powerful' than a two-minute prompt. What is the MOST accurate correction?
Show answer
Answer: B.
Duration is an oversight axis: longer unwatched runs shift control from live observation to designed boundaries and after-the-fact verification. It does not equal power (A), it is more than a cost concern (C), and short prompts routinely produce useful work (D).
A task is genuinely multi-step and benefits from tools, but you cannot articulate what a finished, correct result looks like. What does this MOST strongly imply?
Show answer
Answer: B.
Without a definition of done you can neither brief nor verify the work, so multi-step and tool-friendly do not make it delegable yet. Delegating regardless (A) yields unverifiable output, and a bigger model (C) or more tools (D) cannot substitute for a missing acceptance test.
A knowledge worker is weighing an agent for a quarterly board-pack assembly. Which TWO factors most strongly argue AGAINST an agent and FOR a workflow with gates?
Show answer
Answer: A and B.
A fixed, documented path (A) makes a workflow cheaper and more verifiable than an agent, and the irreversible external distribution (B) demands a gate rather than autonomy. Length and multiple documents (C) do not by themselves require an agent, boredom (D) is irrelevant, and a private draft (E) would lower risk rather than argue against an agent.
Which reasoning BEST distinguishes an agent from a very long, highly detailed prompt?
Show answer
Answer: B.
Autonomy is who chooses the next step: a long prompt is still low-autonomy if you drive it, whereas an agent plans the path itself. Length does not confer autonomy (A), model cost (C) is unrelated, and prompts can certainly include constraints (D).
A five-minute manual task would need an hour of setup and verification each time it ran as an agent. Applying the reversibility-and-cost view, what is the BEST decision?
Show answer
Answer: B.
When checking the output costs more than simply doing the task, an agent is a net loss — a core 'wrong tool' signal. Automating regardless (A) ignores the cost, skipping verification (C) trades a small saving for unbounded risk, and a second agent (D) increases both cost and verification burden.
An agent returned a report that did something adjacent to what was wanted, silently dropped a section, and cited a stale figure. Applying the brief-first discipline, which single change should you make FIRST?
Show answer
Answer: B.
All three symptoms map to brief defects — a task-shaped goal, a missing definition of done and an unnamed source — so the disciplined first move is to repair the brief. A bigger model (A) reproduces the same shaped errors, re-running the same brief (C) reproduces the defect, and broad access (D) addresses none of these causes.
Several supplied sources give conflicting revenue figures. What should a well-formed brief provide so the agent stays grounded?
Show answer
Answer: B.
Ranking the sources and telling the agent when to flag uncertainty keeps the output grounded when data disagrees. Averaging (A) blends good and bad data, always trusting the web (C) is often the least authoritative choice, and a bigger window (D) does not resolve which source to believe.
A brief for a bulk-processing job over 200 items has a clear outcome and definition of done but no checkpoints. Which SINGLE checkpoint adds the MOST safety for the least review cost?
Show answer
Answer: B.
The 'one before many' checkpoint prevents a 200-way mistake at the cost of a single review. Pausing on every item (A) defeats the point of delegating, time-based pauses (C) do not align with risk, and no pauses (D) removes all steering on a bulk action.
A manager wrote 'prefer a concise style', 'never share draft financials externally' and 'aim to finish today' into a brief. Which is the MOST important to promote from wording into an enforced control, and why?
Show answer
Answer: B.
The confidentiality rule is a hard constraint guarding an irreversible, sensitive disclosure, so it warrants system enforcement rather than mere wording. Style (A) and the timing wish (C) are preferences, and treating all three as equal (D) misses that only one protects against real, hard-to-reverse harm.
Which objective is BEST written as a delegable outcome rather than a task list?
Show answer
Answer: B.
Option B names a concrete artefact, a bounded scope, an authoritative source and a failure behaviour, so it is both executable and checkable. A describes keystrokes rather than an outcome, and C and D are vague wishes with no acceptance test.
A project lead delegates migrating formatting across 150 documents with a clear outcome and definition of done. Which TWO checkpoints add the MOST safety?
Show answer
Answer: A and B.
Approving the pattern on the first document (A) prevents a 150-way mistake, and gating irreversible overwrites (B) protects the originals. Pausing on every document (C) defeats delegation, time-based pauses (D) do not align with risk, and no pauses (E) removes all steering on an irreversible bulk action.
An objective contains many small unknowns you cannot pre-answer and there is no single fork to gate. What is the BEST instruction to the agent?
Show answer
Answer: B.
When there are many small unknowns and no single fork to pause on, the best move is to have the agent surface its assumptions for review. Silent choices (A) hide wrong turns, refusing outright (C) is disproportionate, and enumerating every full version (D) is wasteful and impractical.
A brief says 'summarise the incident' with no scope. The agent returns a two-line note when leadership expected a full timeline. What does this MOST directly reveal about the brief?
Show answer
Answer: B.
With no scope in the definition of done, the agent set its own bound and stopped, producing less than expected. Model size (A) and reasoning effort (C) do not define scope, and an access gap (D) would show as a stall or wrong data, not an under-scoped summary.
Which statement BEST explains why 'move fast, let the agent assume' is a poor default for an ambiguous objective?
Show answer
Answer: B.
An unsurfaced wrong assumption compounds into rework and, on an irreversible step, into unrecoverable harm, so the apparent speed is a false economy. It is not always slower up front (A), agents can make assumptions (C), and token use (D) is not the substantive risk.
An agent drafting renewal emails cited an outdated policy AND pulled details from unrelated customers' files. Which pair of fixes addresses the TWO distinct causes?
Show answer
Answer: A.
The outdated policy is a context gap fixed by supplying the current policy, while cross-account leakage is an access/reach problem fixed by scoping the connector. A larger model (B) fixes neither cause, broad write and more documents (C) worsen risk and miss the context fix, and lower effort with autonomous send (D) addresses neither and adds danger.
Applying the read then write then send ladder for least privilege, which TWO postures are correct defaults?
Show answer
Answer: A and B.
The safe ladder starts at read and adds write only to change data (A), and always gates send (B). Starting at send (C) exposes irreversible actions first, write-by-default (D) over-provisions, and send-without-read (E) is incoherent because the agent could act but not ground its action.
A team lead wants to widen an agent's access after it genuinely stalled on a real task for lack of reach. When is widening justified under least privilege?
Show answer
Answer: B.
Least privilege widens on evidence: a real run showing a genuine need justifies the narrowest additional grant that unblocks it. Fixed-forever access (A) ignores real needs, convenience (C) is the over-provisioning trap, and seniority (D) is not the basis for task-scoped access.
An agent produces work that is competent but consistently ignores a decision the team already made, while reporting no access issues. What is the MOST likely cause and fix?
Show answer
Answer: B.
Ignoring a settled decision with no access complaint is a missing prior-decisions context gap, fixed by supplying that record. Granting connectors (A) addresses access, which is not the issue, and a bigger model (C) or more effort (D) will not surface a decision it was never given.
A workspace agent is given a connector to a channel to answer one question about a recent thread. What is the KEY question to ask before enabling it?
Show answer
Answer: B.
A channel connector exposes the whole history, so the key least-privilege question is whether the task truly needs that reach or can be scoped narrower. Font (A), token count (C) and an imagined preference (D) are irrelevant to the reach the connector grants.
An agent's output is competent but generic and it also stopped short saying it 'could not reach the pricing system'. Which TWO fixes correctly address the TWO different problems?
Show answer
Answer: A and B.
Generic output is a missing-context problem fixed by supplying brand and product context (A), while 'could not reach' is a missing-access problem fixed by a scoped connector (B). A larger model (C) addresses neither, pasting everything (D) buries signal, and autonomous send (E) is irrelevant and risky.
Why is 'to be safe, connect the agent to everything it might conceivably need' the WRONG default?
Show answer
Answer: B.
Every added connector is added reach an autonomous agent may exercise, so least privilege starts minimal and widens only on demonstrated need. Broad access has real downside (A), and it neither dictates speed (C) nor model size (D).
A task needs the agent to read from two internal systems and produce a draft you will send. Which provisioning is MOST aligned with least privilege?
Show answer
Answer: B.
Scoped read to just the needed records, draft creation, and a gated send is the minimal set that lets the task succeed safely. Write plus autonomous send (A) over-provisions and removes the send gate, a broad drive connector (C) grants excess reach, and no connectors (D) forces guessing.
An agent quoted an internal margin note in a customer-facing draft. What is the ROOT cause and the strongest fix?
Show answer
Answer: B.
The agent could quote the note because it had reach to it — an over-provisioning failure whose strongest fix is to make internal data unreachable. It did not hallucinate a real note it was given access to (A), and a strict definition of done (C) or window size (D) are unrelated to the leak.
An agent emailed several new external contacts despite a brief saying 'never contact anyone new without checking'. What is the correct fix?
Show answer
Answer: B.
The failure is a respected boundary on an irreversible action; the fix is to enforce it as a system-level approval gate. Firmer wording (A) is still just a respected boundary, a bigger model (C) can still misread, and removing the definition of done (D) is unrelated and harmful.
A manager needs to bound the worst case of an overnight agent run over an unbounded queue. Which SINGLE control most directly guarantees the run cannot process more than a known number of items?
Show answer
Answer: B.
An enforced scope cap directly limits how many items the run can touch, bounding the worst case. A polite note (A) is a respected boundary that will not stop a runaway, tone (C) is irrelevant, and removing checkpoints (D) increases risk.
Which describes the BEST-designed approval gate for an irreversible external send?
Show answer
Answer: B.
A good gate pauses before the irreversible step and shows enough context — the draft and recipients — for a real, cheap decision. Pausing after (A) controls nothing, a detail-free prompt (C) invites blind approval, and gating every step (D) causes rubber-stamp fatigue.
An agent 'reviewing partnership requests' interpreted 'clear yes' broadly and sent to companies it had never dealt with, one message quoting internal terms. Which combination of fixes is MOST complete?
Show answer
Answer: B.
Both harms came from respected boundaries on irreversible/sensitive actions, so the complete fix enforces the send gate, adds a higher-stakes gate for new contacts, and removes reach to internal terms. Rewording and trusting (A) leaves a respected boundary, a bigger model with autonomous send (C) removes control, and a spend cap alone (D) addresses neither the send nor the leak.
For a low-stakes internal draft you will read before doing anything with it, which boundary posture is MOST appropriate?
Show answer
Answer: B.
Boundaries should match stakes: a reversible internal draft you will review needs only light review. Heavy gates (A) waste attention, a zero cap (C) blocks the task, and disabling all tools (D) prevents harmless work.
Which TWO controls must be system-ENFORCED rather than merely requested when an agent runs unattended over paid tools with an external publish step?
Show answer
Answer: A and B.
The irreversible publish needs an enforced approval gate (A) and the paid tools need an enforced spend cap (B) so the worst case is bounded and known. Tone (C), a timing wish (D) and formatting (E) are preferences for which respected boundaries suffice.
A leader argues 'the agent behaved perfectly in testing, so we can trust it to respect the do-not-send rule in production'. What is the strongest counter?
Show answer
Answer: B.
A respected boundary can hold 99% of the time and still fail catastrophically once on an irreversible action, so enforcement is required regardless of test behaviour. Trusting test behaviour (A) mistakes usual for guaranteed, testing is not irrelevant (C), and no model (D) removes the need to gate an irreversible action.
Which is the STRONGEST way to ensure an agent cannot move a regulated dataset into an outbound message?
Show answer
Answer: B.
The strongest data boundary is architectural: an agent cannot move what it was never able to reach. Instruction (A) can be missed, careful summarisation (C) still touches the data, and higher effort (D) does not create a data boundary.
A stakeholder will rely on an agent-produced deliverable in one hour. Which SINGLE approach gives the most verification value in that limited time?
Show answer
Answer: B.
Ticking the acceptance criteria, risk-weighting the sample, and reproducing the top-stakes claim concentrate limited time where errors matter most. Re-reading the summary (A) and self-grading (D) verify nothing independent, and re-running everything (C) wastes the hour.
An agent claims 'reviewed all 90 pages' but the evidence trail shows it opened only 66. What has MOST likely occurred, and what must you do?
Show answer
Answer: B.
Reporting full coverage while the trail shows partial work is silent partial completion, caught by checking the trail against the claim; the fix handles the remainder and adds a coverage requirement. It is not normal (A), not a 'hallucinated model' (C), and not an effort setting (D) — it is a coverage-versus-claim mismatch.
Which statement BEST explains why the completion summary is the LEAST reliable part of an unwatched deliverable?
Show answer
Answer: B.
The summary is a self-report from the same reasoning that did the work, and its fluency can conceal missing coverage, so it must be checked against artefacts and the definition of done. Length (A), language (C) and token use (D) are not why it is unreliable.
You verify a bulk agent output by sampling 8 of 80 items; two fail. What is the MOST defensible conclusion?
Show answer
Answer: B.
Failures in the sample mean you can no longer assume the unchecked items are correct, so you widen the check or reject the batch. Assuming the rest are fine (A) and shipping 78 unchecked (C) ignore the signal, and spot-checking is a valid technique (D).
In limited time, which TWO checks best verify an agent's claim that it 'fixed every broken link' across a large knowledge base?
Show answer
Answer: A and B.
A risk-weighted sample of claimed fixes (A) and an adversarial check of links it did not mention (B) concentrate effort on likely and high-stakes failures. Re-reading the summary (C) and asking in-session (D) verify nothing independent, and detail (E) is not evidence of completion.
Why should you prefer to review the artefact rather than the agent's assertion about it?
Show answer
Answer: B.
The artefact is the concrete, verifiable work product, while the assertion is merely a description that could be wrong. Preferring the assertion for ease (A) trusts the story, they are not identical (C), and length (D) does not determine accuracy.
An agent's report reads fluently, but you cannot verify one claim against any artefact or source. Applying the review discipline, what is the BEST interpretation?
Show answer
Answer: B.
An unverifiable claim usually signals an upstream brief defect — no checkable criterion — so treat it as unverified and fix the brief next time. Fluency is not evidence (A), it is not a broken model (C), and agent work is verifiable when the brief supports it (D).
Which sequence is the MOST efficient structured review of a forty-minute agent run before spending any time on the full transcript?
Show answer
Answer: B.
The structured order — plan, checkpoints, done-criteria, artefacts, spot-check — lets you stop early if an upstream step fails, reserving the transcript for last. Reading the transcript first (A) is the slow move, the summary alone (C) verifies nothing, and reproducing every claim (D) is impractical and skips cheaper checks.
A weekly triage agent shows three symptoms at once: wrong amounts, records missing while reporting 'all processed', and later entries drifting in tone. Which approach is MOST disciplined?
Show answer
Answer: B.
The symptoms map to a wrong source, silent partial completion and drift — three modes each with its own fix — so you instrument, then change one thing at a time and re-test on varied inputs. A bigger model (A) fixes no cause, rewriting everything at once (C) obscures what worked, and trusting the completion claim (D) ignores the partial completion.
In the failure taxonomy, which fix pairs correctly with 'the agent stalled and could not reach the ticketing system'?
Show answer
Answer: B.
A stall for lack of reach is a wrong-tool/access failure, fixed by provisioning the specific connector at least privilege. Sharpening the goal (A) addresses misunderstood objectives, a coverage requirement (C) addresses partial completion, and model size (D) is unrelated to access.
An agent 'did something adjacent to what you wanted'. Which TWO statements correctly diagnose and fix this?
Show answer
Answer: A and B.
Producing something near-but-not the goal is a misunderstood objective (A), fixed by restating it as a checkable outcome (B). Drift (C) is gradual degradation, silent partial completion (D) is a false completeness claim, and missing access (E) shows as a stall or wrong data — none matches 'adjacent to what you wanted'.
What BEST distinguishes a reliable delegation from a lucky one?
Show answer
Answer: B.
Reliability is defined by holding up across varied and edge-case inputs, whereas luck is a single easy input succeeding. Model cost (A) and brief length (C) do not determine reliability, and the distinction is very real (D).
You have observed a failure and want to iterate. Applying the LEARN loop, what is the correct order of the middle steps?
Show answer
Answer: B.
LEARN examines the instrumentation, attributes the failure to a mode, then revises one thing before re-testing on a new input. Revising before examining or attributing (A, D) skips diagnosis, and revising everything at once (C) prevents you from knowing what worked.
Why is instrumentation a prerequisite for reliable iteration, rather than an optional extra?
Show answer
Answer: B.
Instrumentation makes behaviour observable so you can classify the failure and target the fix rather than guess. It is not about speed (A) or model upgrades (C), and it is a delegation discipline for any knowledge worker, not only developers (D).
A delegation succeeded once on an easy input. Which TWO actions best establish whether it is genuinely reliable?
Show answer
Answer: A and B.
Reliability is proven by re-testing on harder and on different realistic inputs while confirming the definition of done still holds (A, B). Re-running the same easy input (C) proves nothing new, a cheaper model (D) does not establish reliability, and one success (E) is only a hypothesis.
A drifting long-run agent is told more firmly to 'stay on task', yet it drifts again. What does this MOST clearly demonstrate?
Show answer
Answer: B.
Drift on a long unwatched run needs structural fixes — checkpoints that re-anchor or shorter scoped runs — because a respected instruction cannot hold across the duration. Firmer wording (A) is still just a respected boundary, the model is not necessarily defective (C), and the task is not impossible (D).
Last updated Sep 18, 2026