Agents and Workflows
Agents – Mock Exam 1
A 50-item, domain-weighted independent mock exam for the Agents and Workflows track, with full explanations and a readiness indicator.
This is a full-length, domain-weighted independent mock exam for the Agents and Workflows track. It is built from publicly available OpenAI learning objectives and is not an official OpenAI assessment. All 50 questions are new and do not repeat the domain-page items. Use this as your diagnostic: sit it first to find your two weakest domains, then revise before mock exam 2.
Instructions
- Time: 60 minutes, matching the length of the interactive sitting.
- Items: 50 multiple-choice and multiple-response questions. Each item states how many answers to select.
- Selection: for a multiple-response item you must select all correct options and no incorrect ones to earn the mark; there is no partial credit.
- No guessing penalty: answer every question — a wrong answer costs nothing beyond the mark.
- Target: aim for at least 80% raw (40 of 50) before you take the real Academy Agents and Workflows assessment. The 80% line matches the Academy badge threshold.
- Work each question before expanding the answer.
Domain distribution
| # | Domain | Weight | Items here |
|---|---|---|---|
| 1 | What an Agent Is and When to Use One | 16% | 8 |
| 2 | Defining Objectives and Tasks | 18% | 9 |
| 3 | Context, Tools and Permissions | 18% | 9 |
| 4 | Boundaries and Guardrails | 16% | 8 |
| 5 | Reviewing and Verifying Agent Work | 16% | 8 |
| 6 | Reliability and Iteration | 16% | 8 |
Total: 8 + 9 + 9 + 8 + 8 + 8 = 50 items.
Readiness interpretation
This is an independent readiness indicator, not an official score.
| Raw score (of 50) | Band | Interpretation |
|---|---|---|
| 45–50 | 90%+ | Strong readiness across all domains |
| 40–44 | 80–89% | Assessment ready; review any weak domain |
| 35–39 | 70–79% | Building confidence; targeted revision advised |
| Below 35 | under 70% | Keep learning; revisit D2 and D3 first |
The 80% band matches the Academy badge threshold, so treat 40 of 50 as your minimum before sitting the real assessment.
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your readiness indicator, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A finance analyst needs a single quarterly figure double-checked against one attached spreadsheet, and will read the answer before doing anything with it. Which mode fits best?
Show answer
Answer: B.
One question answered from one supplied source, read before acting, is a prompt — minimal autonomy, tools and duration. An agent (A, D) adds setup, cost and blast radius for no benefit, and the email in D is an unnecessary irreversible step. A five-step workflow (C) over-engineers a single lookup.
Which characteristic most clearly signals that a task genuinely needs an agent rather than a workflow?
Show answer
Answer: B.
An agent earns its keep when the path cannot be fixed ahead of time, so it must plan as it learns. Emotional weight (A), instruction length (C) and requester seniority (D) do not change whether the steps are knowable in advance, which is what separates a workflow from an agent.
An operations lead schedules an identical month-end reconciliation with the same fixed steps every month and reviews between each step. Which mode is most appropriate?
Show answer
Answer: B.
Fixed, known steps with review points between them are the definition of a workflow, which is cheaper to build and easier to verify. A single prompt (A) cannot carry a sequenced multi-step process, and an agent (C, D) adds autonomy that a fully known path does not require.
Why does a long, unwatched agent run demand different oversight than a short prompt you watch?
Show answer
Answer: B.
Watched work is self-correcting because you can halt a bad step; an unwatched run has already acted by the time you look, so control shifts to boundaries and verification. Cost (A) and model size (C) are not the defining difference, and short prompts can use tools you invoke (D).
A colleague says 'this report is really complicated, so it obviously needs an agent'. What is the soundest reply?
Show answer
Answer: B.
Fit is decided by the four axes, not by how hard a task feels; a complicated task with a fixed path is still a workflow. Treating difficulty as the deciding factor (A) is the trap, blindly splitting (C) ignores whether steps are fixed, and model choice (D) is unrelated to the mode.
A knowledge worker wants to hand off compiling a weekly market-news digest whose sources vary week to week, running unattended for about ten minutes and producing only a draft. Which mode fits and why?
Show answer
Answer: C.
Varying sources mean the path is discovered as it goes, and an unattended multi-step run producing a reversible draft is agent territory. It is not a prompt (A, D) because it is not one watched step, and not a workflow (B) because the steps are not fixed.
Which TWO situations most clearly make an agent the wrong tool for the job?
Show answer
Answer: A and C.
A missing definition of done (A) means you cannot brief or verify the work, and verification costing more than the work (C) makes delegation a net loss — both are classic 'wrong tool' signals. Multi-tool unwatched runs (B) and discovery-driven paths (D) point toward an agent, and reversible drafts (E) simply keep the risk low.
What best captures what 'autonomy' means when distinguishing an agent from a detailed workflow?
Show answer
Answer: B.
Autonomy is about who decides the path: in a workflow you sequence the steps, whereas an agent plans and chooses them. Instruction length (A) measures size not autonomy, and streaming speed (C) and output format (D) are unrelated.
Which delegated objective is most ready to hand to an agent?
Show answer
Answer: C.
Option C states a concrete outcome, a bounded scope and a named source of truth, so the agent can execute and you can verify. A, B and D are wishes with no defined artefact or acceptance test.
Which TWO properties make a definition of done useful for both the agent and your later verification?
Show answer
Answer: A and B.
An observable definition of done (A) lets you verify by ticking criteria, and a failure-aware one (B) surfaces gaps instead of hiding fabrication — both serve execution and verification. Model settings (C), a run time (D) and a connector list (E) are other parts of a brief, none of which is what makes 'done' checkable.
A brief says 'never quote a supplier's confidential pricing' and 'try to keep the memo under one page'. How should the agent treat these two instructions?
Show answer
Answer: B.
A confidentiality rule stated as 'never' is a hard constraint, while a 'try to keep' length target is a preference it can exceed when needed. Treating the confidentiality rule as optional (A, D) is unsafe, and treating the length target as hard (C) could block otherwise good work.
An agent blends this year's headcount with a figure that turns out to be two years old. What is the most likely root cause and fix?
Show answer
Answer: B.
Blending stale and current data is the classic missing-source-of-truth symptom, fixed by naming which source is authoritative and how to resolve conflicts. Reasoning effort (A) governs how hard it thinks, model size (C) and mere length (D) do not tell it which figure to trust.
A brief asks an agent to apply a new naming convention across 80 files. Where does the single highest-value checkpoint usually belong?
Show answer
Answer: B.
The 'one before many' checkpoint converts an 80-way mistake into a one-way one by approving the approach before it scales. Reviewing only at the end (A) means the error already repeated, time-based pauses (C) do not align with risk, and no checkpoints (D) remove your steering.
An objective is ambiguous on one specific, identifiable point — which of two teams the report is 'for'. What is the best instruction to the agent?
Show answer
Answer: B.
A single identifiable fork is best handled by a checkpoint: pause and ask before the assumption propagates. Deciding silently (A) risks a wrong path, skipping it (C) leaves the task incomplete, and producing two full versions (D) is wasteful when one question resolves it.
An agent could not find a figure it needed, so it produced a plausible-looking number to complete the table. What should the brief have instructed?
Show answer
Answer: B.
A failure-aware definition of done tells the agent to surface gaps instead of inventing data. Filling gaps (A) hides fabrication, halting the entire task (C) is disproportionate when one gap can be flagged, and model size (D) does not change fabrication behaviour.
In the five-mode failure taxonomy, which mode is described as the most dangerous because the agent reports success while having done only part of the work?
Show answer
Answer: B.
Silent partial completion is the most dangerous mode because nothing signals a problem — the agent claims full coverage while some work is missing. Misunderstood objective (A) produces adjacent work, missing context (C) produces off-brand or wrong-fact output, and drift (D) is gradual degradation over a long run — each is visible in a different way.
Which TWO items belong in a strong definition of done for a slide-ready competitor summary?
Show answer
Answer: A and C.
Coverage of all four competitors (A) and a size bound of one slide with six bullets (C) are observable acceptance criteria you can tick. Model choice (B) and run time (D) are operational, and tone (E) is a preference rather than an acceptance test for a data summary.
Why is 'state the outcome, not the keystrokes' the recommended way to write a goal?
Show answer
Answer: B.
An outcome lets the agent plan the path while remaining verifiable, whereas a list of your keystrokes constrains it and obscures what 'done' means. Length (A) is not the point, outcomes do not require larger models (C), and agents can follow steps (D) — they simply plan better from an outcome.
An agent returns competent but generic, off-brand copy. Which fix is most appropriate?
Show answer
Answer: B.
Off-brand output is a missing-context symptom: the agent lacks the company knowledge that defines the brand. More tools (A) and write access (D) add access rather than context, and a larger model (C) still will not know your brand without being told.
A task only requires the agent to answer questions from one uploaded policy document. What access should it receive?
Show answer
Answer: B.
Least privilege means granting exactly the read access the task needs. Web and write (A), broad connectors (C) and send (D) all add blast radius the task never requires.
Which set best names the four kinds of context an agent typically needs?
Show answer
Answer: A.
The four context types are the task brief, the source documents for the task, relevant company knowledge and prior decisions already settled. B lists sampling settings, C lists account facts and D lists formatting — none is the context an agent reasons from.
You connect an agent to an entire shared drive so it can open one specific file. What is the risk?
Show answer
Answer: B.
A connector exposes everything it reaches, and an autonomous agent may use any of it toward the goal, surfacing data far beyond your intended file. It will not politely restrict itself (A), and speed (C) and model choice (D) are unaffected by scope.
An agent reports 'I cannot access the CRM' and stops. What is the correct fix?
Show answer
Answer: B.
'Cannot access' is a missing-access problem, fixed by granting the specific connector scoped to the records required. Pasting documents (A) addresses context, and model size (C) or reasoning effort (D) do not grant access.
Why can supplying too much context actually hurt an agent's output?
Show answer
Answer: B.
Dumping everything in buries the authoritative signal and can make the agent over-weight irrelevant or outdated material. More is not always better (A), and it neither changes pricing (C) nor disables tools (D).
A task must produce an email draft you will send yourself after review. How should access be arranged for least privilege?
Show answer
Answer: B.
Draft-before-send with the irreversible send gated behind approval is the least-privilege, safe arrangement. Autonomous send (A) removes the gate, whole-system write (C) is far broader than needed, and no tools (D) under-provisions a task that must produce a draft.
An agent keeps reopening a question the team settled last quarter. Which kind of context was missing?
Show answer
Answer: A.
Relitigating a settled question is the signature of missing 'prior decisions' context; supply the record of what was decided. A bigger window (B) will not help if the decision was never provided, and web (C) and write (D) are access, not the missing context.
Which TWO moves best reduce an agent's blast radius without reducing its ability to do a read-only research task?
Show answer
Answer: A and B.
Read-instead-of-write (A) and scoping a connector to the needed folder (B) shrink blast radius while leaving a read-only task fully doable. Every connector (C) and autonomous send (D) enlarge blast radius, and removing the definition of done (E) harms the task without improving safety.
An agent worked perfectly on the one clean input tested but failed on the real batch. Which upstream lesson does this teach about proving a delegation?
Show answer
Answer: B.
The clean input was the lucky case; reliability is only established by re-testing on varied and edge-case inputs. Treating one success as proof (A) is exactly the trap, model size (C) rarely explains a brief that was never stress-tested, and batch size alone (D) does not make delegation impossible.
Which action most requires an enforced approval gate before it can fire?
Show answer
Answer: B.
Posting to an external customer channel is irreversible, so it needs an enforced gate before it fires. Drafting (A), analysing (C) and rewriting (D) are reversible and safe with light review because nothing has left the workspace.
What is the key difference between a respected boundary and an enforced boundary?
Show answer
Answer: B.
A respected boundary is instruction the agent usually honours, whereas an enforced boundary is a system control that removes the possibility of crossing. Tone (A) is irrelevant, enforcement is not model-specific (C), and the difference is central (D).
For an irreversible action, why is 'just instruct the agent clearly not to do it' insufficient?
Show answer
Answer: B.
Instruction depends on correct understanding, and an autonomous agent can cross a requested line by misreading the goal — unacceptable when the action is irreversible. Clear instructions do not always hold (A), token cost (C) is irrelevant, and agents do not ignore all instructions (D) — they can misinterpret them.
An agent that can iterate and call paid tools is told 'please don't spend too much'. What is the problem?
Show answer
Answer: B.
A vague requested limit will not halt a runaway; an enforced spend cap makes the worst case bounded and known. The instruction alone is not enough (A), loops can occur with paid tools (C), and an enforced cap stops a runaway rather than slowing the model (D).
What is the STRONGEST data boundary against an agent leaking a sensitive internal document?
Show answer
Answer: B.
The strongest data boundary is architectural: an agent cannot leak what it was never able to reach. Instruction (A) can be missed, careful summarisation (C) still touches the data, and model size (D) creates no data boundary.
Why can gating every single action be as harmful as gating none?
Show answer
Answer: B.
Over-gating produces rubber-stamp fatigue, defeating the gate when it matters. More gates are not always better (A), and gates neither disable tools (C) nor change the model (D).
Which TWO boundaries should be system-enforced rather than merely written in the brief?
Show answer
Answer: B and C.
An approval gate on irreversible sends (B) and a hard spend cap on paid-tool use (C) protect against real, hard-to-reverse harm and must be enforced. Style (A), tone (D) and a soft timing wish (E) are preferences for which respected boundaries are fine.
An agent 'improving' a document overwrote the original, losing the prior version. Which boundary would have prevented this?
Show answer
Answer: B.
Overwriting can be irreversible, so gating overwrites or forcing versioned saves preserves the original. Tone (A), context window (C) and web search (D) do nothing to protect the file.
An agent ends a run with 'I completed every task successfully.' What is the best next step?
Show answer
Answer: B.
The summary is a claim; verification means checking artefacts and acceptance criteria independently. Trusting the summary (A) is the fluency trap, a same-session self-check (C) reuses the agent's own reasoning, and re-running everything (D) wastes the delegation without verifying the original.
Why is 'ask the agent whether it did the work correctly' a weak verification method?
Show answer
Answer: B.
Asking the agent to self-assess in the same session carries the original reasoning and bias, so it may confirm a mistake. Token cost (A) is not the issue, agents can describe their work (C) — that is the unreliable claim — and it is not the strongest method (D).
In verifying agent work, what is an 'artefact'?
Show answer
Answer: B.
An artefact is the actual work product you can inspect and check, unlike the agent's description of it. The summary (A) is an assertion, and effort setting (C) and start time (D) are run metadata, not the work product.
You spot-check 6 of 60 bulk items and one fails. What should you conclude?
Show answer
Answer: B.
A failure in the sample means you can no longer assume the unchecked items are correct, so you widen the check or reject. Assuming the rest are fine (A) ignores the signal, fixing only one (C) leaves 53 unverified, and spot-checks are useful, not unreliable (D).
What is the most efficient way to review a long, mostly-unwatched agent run?
Show answer
Answer: B.
Structured review — plan, checkpoints, done-criteria, artefacts, spot-check — is faster and catches more than a linear read. The full transcript (A) is a last resort, trusting the summary (C) verifies nothing, and shortening the transcript (D) does not verify the work.
For a single high-stakes number in an agent's report, what is the strongest check?
Show answer
Answer: B.
Independently reproducing the critical claim confirms it without relying on the agent's own reasoning. Shown working (A) can still be wrong, a same-session recompute (C) reuses the original reasoning, and rounding (D) conceals rather than checks.
Which TWO checks best verify an agent's claim that it 'updated every out-of-date figure' in a document?
Show answer
Answer: A and B.
Sampling the claimed updates (A) checks the reported work, and checking figures it did not mention (B) is the adversarial check that catches what the summary hides. Re-reading the summary (C) and asking in-session (D) verify nothing independent, and detail (E) is not evidence of completion.
You find you cannot verify one of an agent's claims at all against any evidence. What does this most likely indicate?
Show answer
Answer: B.
An unverifiable claim usually reflects an upstream brief defect — no checkable criterion — which you fix in the next iteration. Inability to verify does not make the claim correct (A), it is not a broken model (C), and agent work is verifiable when the brief supports it (D).
An agent reports 'all 40 records processed' but 6 are missing from the output. Which TWO responses are correct?
Show answer
Answer: A and B.
Reporting full completion while work is missing is silent partial completion (A), and the reliability fix is a coverage requirement that forces the agent to flag what it could not do (B). Trusting the claim (C) ignores the mismatch, a bigger model (D) does not address a missing coverage check, and removing the definition of done (E) removes your only means of catching it.
An agent starts a long run well, then later entries wander off task and over-elaborate. What is this and the correct fix?
Show answer
Answer: B.
Gradual wandering on a long unwatched run is drift; the structural fix is checkpoints that re-anchor plus shorter scoped runs. More documents (A), more tools (C) and a bigger model (D) do not address a duration-driven loss of focus.
A delegation succeeded on the one input you tried. What should you conclude?
Show answer
Answer: B.
A single success is a hypothesis; reliability is demonstrated across varied and edge-case inputs. Declaring it reliable (A) risks a fragile brief, re-running the same input (C) proves nothing new, and cost choices (D) do not establish reliability.
Why change only one thing at a time when iterating a delegation brief?
Show answer
Answer: B.
Changing one thing isolates cause and effect, so you learn what works and build reliability deliberately. It is not about slowing down (A) or token use (D), and agents read full briefs, not one instruction (C).
A delegation fails and a colleague immediately suggests switching to the most capable model. What is the best response?
Show answer
Answer: B.
The disciplined move is to classify the failure and fix its upstream cause; the model is the last lever because it rarely addresses a vague brief or missing context. Model-first thinking (A, C) skips diagnosis, and abandoning (D) gives up before diagnosing.
Which TWO are structural fixes for drift over a long unwatched run?
Show answer
Answer: A and B.
Checkpoints (A) and shorter scoped runs (B) are structural controls that stop drift accumulating. 'Stay focused' (C) is a respected boundary that will not hold on a long run, higher effort (D) does not address duration-driven wandering, and removing the definition of done (E) removes your coverage check.
Last updated Sep 18, 2026