AI Cert Prep
Type to search documentation.

Applied AI Foundations

Applied AI Foundations · Mock Exam 1

A 50-item independent mock exam for the Applied AI Foundations track, weighted to the six domains, with full explanations and a readiness indicator.

This is Mock Exam 1 for the Applied AI Foundations track: 50 items across the six domains. It is an independent mock exam built from publicly available OpenAI learning objectives — not an official OpenAI assessment, not official exam questions, and not endorsed by OpenAI. Use it as your diagnostic: sit it first to find your two weakest domains, then study those domain pages before you attempt Mock Exam 2. For how the real credential works, see the assessment model.

Instructions

  • Time: 60 minutes, matching the length of an Applied AI Foundations study session; you can also run it untimed on a first diagnostic pass.
  • Items: 50, single-response and multiple-response. Each item states how many answers to select.
  • Selection: for a select-one item choose exactly one option; for a select-two item you must choose both correct options and no incorrect one to score the item.
  • No guessing penalty: answer every item — an unanswered item scores the same as a wrong one.
  • Target: aim for ≥ 80% raw (≈ 40/50) before you take the real OpenAI Academy assessment, which itself passes at 80%.

Domain distribution

#DomainWeightItems here
1Finding and Scoping Opportunities16%8
2Decomposing Work into Steps18%9
3Inputs, Outputs and Contracts16%8
4Choosing the Right Capability20%10
5Review Points and Human Oversight16%8
6Repeatability and Improvement14%7

Total: 8 + 9 + 8 + 10 + 8 + 7 = 50 items.

Readiness interpretation

This is an independent readiness indicator, never an official score.

Raw score (of 50)BandInterpretation
40+ (80%+)Assessment readyYou are tracking the 80% Academy badge threshold; take Mock Exam 2 to confirm
35–39 (70–79%)Building confidenceClose; revise your weakest one or two domains and re-test
under 70%Keep learningWork through the domain pages before re-attempting
45+ (90%+)Strong readinessComfortable margin across the domains

Take the mock exam

Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · Finding and Scoping OpportunitiesSelect one

    A team lead notices that onboarding-email personalisation happens about 30 times a week, takes 12 minutes each, tolerates small wording errors, and uses only public marketing copy. On the FTED screen, how should this task be classified?

    • A. A poor candidate because 12 minutes is trivial
    • B. A strong candidate: high frequency, real time cost, correctable errors, low-sensitivity data
    • C. A poor candidate because email is always high-risk
    • D. Undecidable without knowing the model price first
    Show answer

    Answer: B.

    Both value axes (frequency × time = 6 hours/week) and both safety axes (correctable errors, public data) score well, so it sits in the automate-now zone. Twelve minutes across 30 instances is not trivial (A). Marketing email is not inherently high-risk (C). Model price is a build detail that comes after the screen, not a blocker to classification (D).

  2. Q2D1 · Finding and Scoping OpportunitiesSelect one

    Which statement best captures why importance is not the same as value when screening automation candidates?

    • A. Important tasks are always more expensive to automate
    • B. Value comes from frequency times time cost, whereas importance is about visibility and stakes, so a rare-but-important task offers little payback for a workflow
    • C. Importance and value are actually identical in the FTED model
    • D. Important tasks can never touch sensitive data
    Show answer

    Answer: B.

    A once-a-year showpiece may be highly important yet has almost no repeat frequency, so the workflow you build never amortises. Cost of building (A) is not what distinguishes the two concepts. They are explicitly not identical (C). Sensitivity is a separate axis unrelated to importance (D).

  3. Q3D1 · Finding and Scoping OpportunitiesSelect one

    A workflow would draft press statements that are published verbatim to the company newsroom. Which single FTED axis most constrains how you build it?

    • A. Frequency, because press statements are rare
    • B. Time cost, because they are quick to write
    • C. Error tolerance, because a wrong external statement is hard to walk back and is externally visible
    • D. None of the axes apply to writing tasks
    Show answer

    Answer: C.

    Low error tolerance on an external, hard-to-retract artefact forces a mandatory human gate before publication. Frequency (A) affects payback, not how carefully you gate. These statements are not necessarily quick (B). FTED applies to any recurring task including writing (D).

  4. Q4D1 · Finding and Scoping OpportunitiesSelect one

    You must pick one pilot. Task X saves 8 hours/week with a light weekly sample; task Y saves 9 hours/week but needs a human to approve every one of its irreversible outputs. Which do you build first and why?

    • A. Y, because raw hours saved are higher
    • B. X, because after subtracting the far smaller review burden its net payback is higher and its risk is lighter
    • C. Neither, because both are too risky
    • D. Build both simultaneously to save time
    Show answer

    Answer: B.

    Ranking uses risk-adjusted payback: Y's per-item approval of irreversible output erodes most of its 9 hours and carries higher risk, so X's net saving wins. Raw hours (A) ignore the review cost. X is clearly worth automating (C). Building both at once (D) contradicts the one-pilot constraint and dilutes focus.

  5. Q5D1 · Finding and Scoping OpportunitiesSelect one

    A manager asks you to 'use AI for our whole procurement process'. What is the FIRST scoping move?

    • A. Buy the most capable model available
    • B. Narrow it to one bounded task with a trigger, inputs, output, definition of done, and an out-of-scope boundary
    • C. Build an autonomous procurement agent immediately
    • D. Feed all historical purchase orders into one prompt
    Show answer

    Answer: B.

    'The whole process' is a vague wish; scoping narrows it to a bounded problem you can actually build and test. Buying a model (A) is premature. An autonomous agent (C) skips scoping entirely. One giant prompt (D) is both premature and a decomposition anti-pattern.

  6. Q6D1 · Finding and Scoping OpportunitiesSelect one

    Which element of a scoped opportunity most directly prevents scope creep during the build?

    • A. The trigger
    • B. The list of knowledge files
    • C. The explicit out-of-scope boundary stating what the workflow must not do
    • D. The choice of a larger model
    Show answer

    Answer: C.

    The boundary is what stops 'draft the reply' quietly growing into 'and send it' and 'and issue the refund'. The trigger (A) defines when it runs, not what it must avoid. Knowledge files (B) and model size (D) do not constrain scope.

  7. Q7D1 · Finding and Scoping OpportunitiesSelect two

    Which TWO tasks are the strongest automation candidates on the FTED screen?

    • A. Categorising 300 routine internal support tags per week, corrected at a weekly audit
    • B. Approving loan disbursements over $100,000
    • C. Drafting weekly internal newsletter summaries from public team updates
    • D. Writing the once-a-year corporate sustainability report
    • E. Providing individualised legal advice to customers
    Show answer

    Answer: A and C.

    A and C are frequent, time-consuming, correctable and low-sensitivity — the automate-now profile. Loan disbursements (B) are irreversible and financial. The annual report (D) fails on frequency. Individual legal advice (E) is regulated and high-harm.

  8. Q8D1 · Finding and Scoping OpportunitiesSelect one

    A task passes the value screen easily but touches regulated patient data that policy forbids leaving a sanctioned environment. What does the data-sensitivity axis tell you?

    • A. Lower the value score to compensate
    • B. Nothing, because summarising data is always safe
    • C. Data sensitivity is scored independently and can constrain or block the build regardless of value; use only a sanctioned environment or keep it manual
    • D. Use a cheaper model to reduce the risk
    Show answer

    Answer: C.

    Sensitivity is a separate axis that can veto a high-value task; the answer is a sanctioned workspace with clearance, or manual handling. Sensitivity does not discount value (A). Summarising is not inherently safe (B). Model price (D) has nothing to do with data governance.

  9. Q9D2 · Decomposing Work into StepsSelect one

    A single long prompt produces a report that 'works in the demo but occasionally drops a whole section in production and is hard to debug'. What is the BEST fix?

    • A. Add sentences insisting every section appears
    • B. Decompose it into named single-job steps so each section is an inspectable step that visibly ran or did not
    • C. Switch to the most capable model
    • D. Lower the temperature to zero
    Show answer

    Answer: B.

    Dropped sections and undebuggable output are the mega-prompt signature; decomposition exposes each section as a step. Padding the prompt (A) worsens the blob. A bigger model (C) hides failures rather than exposing them. Temperature (D) does not address structure.

  10. Q10D2 · Decomposing Work into StepsSelect one

    You must summarise fifteen independent regional reports into one national brief that reconciles totals and terminology. Which pattern fits and what is the critical step?

    • A. Sequential; the first summary is critical
    • B. Fan-out/fan-in; the fan-in reconciliation that totals and dedupes is critical
    • C. Draft-then-critique; the draft is critical
    • D. Extract-then-transform; extraction is critical
    Show answer

    Answer: B.

    Independent pieces combined into one consistent output is fan-out/fan-in, and consistency is enforced at the fan-in, which must reconcile rather than concatenate. Sequential (A) misfits independent work. The other two (C, D) are the wrong top-level shape here.

  11. Q11D2 · Decomposing Work into StepsSelect one

    Why is a separate critique step generally better than asking the model to check its own answer in the same reply?

    • A. It always uses fewer tokens
    • B. A same-pass self-review tends to defend the draft, while a separate pass gives a fresh-eyes review, ideally against an explicit checklist
    • C. It removes the need for any human review
    • D. It is always faster
    Show answer

    Answer: B.

    Reviewing in the same pass carries the reasoning that produced the draft; a separate, checklist-driven critique is more objective. It is not about token count (A) or speed (D), and it does not remove human review (C) — often the critique step is where the gate lives.

  12. Q12D2 · Decomposing Work into StepsSelect one

    A workflow reads receipts and outputs a reimbursement total, but a wrong total could be a misread amount OR a misapplied policy rule. Which decomposition makes the failure traceable?

    • A. Draft-then-critique on the prose
    • B. Extract-then-transform: extract structured fields first, then apply rules to the structure
    • C. One more detailed single prompt
    • D. Fan-out/fan-in across receipts
    Show answer

    Answer: B.

    Separating extraction from the rule-based transform makes each independently checkable, so a wrong total traces to reading or logic. Draft-then-critique (A) reviews prose, not this. A detailed single prompt (C) keeps errors blended. Fan-out (D) parallelises but still blends extract and transform per receipt.

  13. Q13D2 · Decomposing Work into StepsSelect one

    When should you STOP splitting a task into more steps?

    • A. Never — more steps are always more robust
    • B. When each step already has one job and a further split adds a handoff but no new inspection point
    • C. After exactly five steps
    • D. When the prompt starts to look long
    Show answer

    Answer: B.

    Splitting earns its keep only when it separates distinct jobs or adds a needed check; beyond that it adds latency and context loss. 'More is always better' (A) and a fixed count (C) are wrong, and prompt length (D) is not the criterion.

  14. Q14D2 · Decomposing Work into StepsSelect one

    A proposal workflow keeps producing an executive summary that describes an earlier version of the body. What ordering fixes this?

    • A. Write the executive summary first to set direction
    • B. Write the executive summary last, generated from the finished body
    • C. Write the summary and body in one prompt
    • D. Skip the summary entirely
    Show answer

    Answer: B.

    A summary must describe finished content, so it is written last in the sequence. Writing it first (A) guarantees drift as the body changes. One prompt (C) is the mega-prompt trap. Skipping it (D) fails the requirement.

  15. Q15D2 · Decomposing Work into StepsSelect one

    A launch-pack workflow produces five assets that sometimes contradict each other on price and key claims. What is the ROOT-CAUSE fix?

    • A. Tell the model to keep prices consistent
    • B. Generate every asset from a single extracted fact sheet so all assets share one source of truth
    • C. Generate the assets twice and keep the longer set
    • D. Use a bigger model
    Show answer

    Answer: B.

    Contradictions come from each asset inventing its own facts; a shared extracted fact sheet gives one source of truth all assets draw from. An instruction to be consistent (A) is padding. Generating twice (C) doubles cost without fixing the cause. A bigger model (D) does not create a shared source of truth.

  16. Q16D2 · Decomposing Work into StepsSelect one

    Which task is best kept as a single step rather than decomposed?

    • A. Producing a five-section report with pricing, timeline and an executive summary
    • B. Classifying a short comment as positive, neutral or negative
    • C. Summarising twelve documents into one consistent brief
    • D. Drafting a policy then critiquing it against a checklist
    Show answer

    Answer: B.

    A single small job with one output does not benefit from splitting; decomposition would only add overhead. The multi-section report (A), the twelve-document brief (C) and the draft-then-critique task (D) all have distinct sub-goals that fail separately and warrant decomposition.

  17. Q17D2 · Decomposing Work into StepsSelect two

    Which TWO signals in a task description point toward the fan-out/fan-in pattern?

    • A. Several independent pieces are processed and then combined into one result
    • B. The combined output must be consistent across all the branches
    • C. You must act on facts buried inside a single document
    • D. A fresh-eyes review pass is the main quality lever
    • E. Each step depends strictly on the previous step's output
    Show answer

    Answer: A and B.

    Independent pieces combined with a consistency requirement is the fan-out/fan-in signature, where the fan-in reconciles. Acting on facts in one document (C) points to extract-then-transform. A review pass (D) points to draft-then-critique. Strict dependence (E) points to sequential.

  18. Q18D3 · Inputs, Outputs and ContractsSelect one

    A step 'sometimes returns a paragraph, sometimes bullets' and a colleague gets a different format than you. What is the ROOT problem?

    • A. The model is too small
    • B. The output shape is not pinned in the step's contract
    • C. The temperature is too high
    • D. The colleague used the wrong account
    Show answer

    Answer: B.

    Inconsistent format across inputs and users means the contract never specified an output shape; pinning a schema, template or fixed fields fixes it. Model size (A) and temperature (C) do not define structure, and the account (D) is irrelevant to format.

  19. Q19D3 · Inputs, Outputs and ContractsSelect one

    Which set best defines the five parts of a step contract?

    • A. Model, temperature, max tokens, stop sequence
    • B. Named inputs, required vs optional, output shape, acceptance criteria, missing-input behaviour
    • C. Prompt, response, cost, latency
    • D. Trigger, owner, deadline, budget
    Show answer

    Answer: B.

    A step contract pins what goes in, what must be present, what comes out, what good means, and what happens on missing input. Option A lists model settings, C lists observability metrics, and D lists project-management fields, none of which is a contract.

  20. Q20D3 · Inputs, Outputs and ContractsSelect one

    A categorisation step runs and its required category list is missing. What should it do?

    • A. Invent plausible categories to keep going
    • B. Fail loudly with a clear reason, because it cannot categorise safely without the list
    • C. Return an empty string silently
    • D. Pick the first category alphabetically
    Show answer

    Answer: B.

    With no safe default and correctness depending on the list, the step must fail loudly rather than guess. Inventing categories (A) fabricates undetectable data, a silent empty string (C) hides the failure, and an arbitrary pick (D) is a disguised guess.

  21. Q21D3 · Inputs, Outputs and ContractsSelect one

    Which is a strong, observable acceptance criterion for a support reply?

    • A. The reply is high quality
    • B. The reply reads professionally
    • C. Answers every question asked, uses no term from the banned list, and ends with a concrete next step
    • D. The reply is comprehensive
    Show answer

    Answer: C.

    Acceptance criteria must be checkable conditions a person or step can verify. 'High quality' (A), 'reads professionally' (B) and 'comprehensive' (D) are all subjective and unverifiable.

  22. Q22D3 · Inputs, Outputs and ContractsSelect one

    You want a step's output parsed by a downstream system. What should the contract require?

    • A. A friendly explanation followed by the data
    • B. Only the JSON, with specified field names, types and allowed values, and no surrounding prose
    • C. A narrative paragraph
    • D. Whatever format the model prefers
    Show answer

    Answer: B.

    Machine consumption needs a strict schema and only the JSON so parsing is reliable. Prose around the data (A) and a narrative (C) break parsers, and letting the model choose (D) reintroduces drift.

  23. Q23D3 · Inputs, Outputs and ContractsSelect one

    After you edit step 2, the workflow breaks at step 3. What is the MOST likely cause?

    • A. Step 3 needs a bigger model now
    • B. Step 2's output no longer matches step 3's required input — a handoff/contract mismatch
    • C. The temperature drifted
    • D. Step 1 is broken
    Show answer

    Answer: B.

    Editing a step commonly changes its output shape, breaking the consumer that relied on the old contract; re-check both edges. Model size (A) and temperature (C) are not implied, and step 1 (D) was not touched.

  24. Q24D3 · Inputs, Outputs and ContractsSelect one

    When is a marked DEFAULT the right missing-input strategy?

    • A. When correctness fully depends on the missing value
    • B. When a safe, documented default exists and you flag that the value was defaulted
    • C. Whenever you want to avoid any interruption
    • D. Never — always fail
    Show answer

    Answer: B.

    A default is appropriate only when it is safe and documented, and it must be marked so review sees the value was missing. If correctness depends on the value (A), fail instead; avoiding interruptions at any cost (C) leads to silent guessing; and 'always fail' (D) is too rigid when a safe default exists.

  25. Q25D3 · Inputs, Outputs and ContractsSelect two

    Which TWO contract rows are most often left blank in flaky workflows?

    • A. The behaviour on missing or malformed input
    • B. The model's release date
    • C. The handoff match to the next step's required input
    • D. The prompt's word count
    • E. The output's colour
    Show answer

    Answer: A and C.

    Undefined missing-input behaviour and an unchecked handoff are the classic gaps that make workflows flaky. Release date (B), word count (D) and colour (E) are not contract rows.

  26. Q26D4 · Choosing the Right CapabilitySelect one

    You repeatedly run the same task and it needs the same style guide and reference files every time. Which rung fits BEST?

    • A. A one-off prompt with the files re-pasted each time
    • B. A Project with the files as knowledge and standing instructions
    • C. An API application
    • D. A workspace agent
    Show answer

    Answer: B.

    Recurring context that must be present every time is the defining signal for a Project. Re-pasting (A) is the waste a Project removes, and an API app (C) or agent (D) over-reaches for a solo recurring task.

  27. Q27D4 · Choosing the Right CapabilitySelect one

    The decisive question separating a Project from a custom GPT is:

    • A. Which model it uses
    • B. Whether other people need to run the task themselves without you
    • C. How many knowledge files it holds
    • D. Its temperature setting
    Show answer

    Answer: B.

    A Project is your recurring workspace; a custom GPT packages the task so others run it independently. Model (A), file count (C) and temperature (D) do not determine which of the two you need.

  28. Q28D4 · Choosing the Right CapabilitySelect one

    A task must run inside your company's web app, on demand, wired to your database. Which capability is required?

    • A. A saved instruction
    • B. A Project
    • C. A custom GPT
    • D. An API application
    Show answer

    Answer: D.

    Embedding in another product, on demand, integrated with a database is exactly what an API application is for. Saved instructions (A), Projects (B) and custom GPTs (C) all live inside ChatGPT and cannot embed programmatically into your app.

  29. Q29D4 · Choosing the Right CapabilitySelect one

    A user needs the latest news about a named company from this week. Which in-conversation capability fits BEST?

    • A. Deep research
    • B. Search
    • C. Data analysis
    • D. Canvas
    Show answer

    Answer: B.

    A quick current-facts lookup is what search is for. Deep research (A) is heavier, for sourced multi-source reports; data analysis (C) is for computation over files; and Canvas (D) is for iterating on documents.

  30. Q30D4 · Choosing the Right CapabilitySelect one

    The task is 'produce a thoroughly sourced comparison of the five leading vendors, with citations'. Which capability fits BEST?

    • A. Plain chat with no tools
    • B. Search for one headline
    • C. Deep research
    • D. Canvas only
    Show answer

    Answer: C.

    A multi-source, cited, comprehensive comparison is the deep-research use case. Plain chat (A) cannot source it reliably, a single search headline (B) under-delivers, and Canvas (D) is an editing surface, not a research tool.

  31. Q31D4 · Choosing the Right CapabilitySelect one

    You uploaded a sales spreadsheet and need outliers found and a trend plotted. What must be switched on?

    • A. File uploads alone
    • B. Data analysis
    • C. Search
    • D. A custom GPT
    Show answer

    Answer: B.

    Computation, outlier detection and plotting over a file require data analysis; uploads alone only let the model read the file (A). Search (C) is for current facts, and a custom GPT (D) is a packaging rung, not a compute capability.

  32. Q32D4 · Choosing the Right CapabilitySelect one

    Ten colleagues need to run the same configured assistant themselves, getting consistent results without your involvement. Which rung?

    • A. A Project you keep to yourself
    • B. A custom GPT shared across the workspace
    • C. A one-off prompt you send them
    • D. An API application
    Show answer

    Answer: B.

    Others running the same task independently with consistent behaviour is the custom-GPT signal. A solo Project (A) does not hand the task to others, a one-off prompt (C) yields inconsistency, and an API app (D) over-reaches for in-ChatGPT self-serve.

  33. Q33D4 · Choosing the Right CapabilitySelect two

    Which TWO statements correctly justify reaching for the lightest capability rung that meets the need?

    • A. Each heavier rung adds build, maintenance and governance cost that only a real need justifies
    • B. Heavier rungs are always slower to run
    • C. Under-reaching wastes time and yields inconsistent results, so you match the rung to the need actually present
    • D. Lighter rungs are always more accurate than heavier ones
    • E. OpenAI policy mandates the lightest rung
    Show answer

    Answer: A and C.

    Climbing the ladder adds total cost of ownership, and under-reaching also has a cost, so you right-size to the actual need. Heavier rungs are not inherently slower (B) or less accurate at lighter rungs (D), and it is a design principle, not a policy rule (E).

  34. Q34D4 · Choosing the Right CapabilitySelect one

    A recurring deliverable is a document you refine over several turns in the same session. Which in-conversation capability best supports the iteration?

    • A. Search
    • B. Deep research
    • C. Canvas
    • D. Data analysis
    Show answer

    Answer: C.

    Canvas is the dedicated surface for iterating on a document or code across turns. Search (A) and deep research (B) gather information, and data analysis (D) computes over files — none is an editing surface.

  35. Q35D4 · Choosing the Right CapabilitySelect one

    A multi-step task can run semi-autonomously with review at checkpoints, executing several actions rather than answering a single question. Which rung fits?

    • A. A saved instruction
    • B. A workspace agent with oversight at checkpoints
    • C. A one-off prompt
    • D. Deep research
    Show answer

    Answer: B.

    Delegated multi-step execution under oversight is exactly a workspace agent. A saved instruction (A) and a one-off prompt (C) answer, they do not execute steps, and deep research (D) is a research capability, not an execution rung.

  36. Q36D6 · Repeatability and ImprovementSelect two

    Which TWO practices make it possible to attribute a quality change to a specific prompt edit and roll it back if it hurts?

    • A. Change only one variable per version
    • B. Use the most capable model available
    • C. Keep the previous version so a regression can be reverted
    • D. Write a longer prompt each time
    • E. Review every single output
    Show answer

    Answer: A and C.

    Changing one variable per version makes effects attributable, and keeping the prior version enables rollback. A bigger model (B) and a longer prompt (D) do not aid traceability, and reviewing every output (E) catches errors but does not attribute which edit caused a change.

  37. Q37D5 · Review Points and Human OversightSelect one

    A workflow sends customer emails automatically with no human check because 'the model is reliable'. What is the problem?

    • A. Nothing — reliability justifies auto-send
    • B. Sending is irreversible and external, a mandatory review trigger, so a human gate is required before send
    • C. The model should be bigger
    • D. The temperature is too high
    Show answer

    Answer: B.

    External, irreversible actions require a human gate regardless of model reliability. 'Reliability justifies it' (A) ignores the mandatory trigger, and model size (C) or temperature (D) do not address the missing gate.

  38. Q38D5 · Review Points and Human OversightSelect one

    A step runs 3,000 times a day, errors are correctable, and stakes per item are low. Which review mode fits BEST?

    • A. Gate every item
    • B. Sampling a tracked fraction, oversampling risky slices
    • C. No review at all
    • D. Deep research on each item
    Show answer

    Answer: B.

    High-volume, correctable, low-stakes work is the textbook case for sampling with tracking. Gating all 3,000 (A) throttles and breeds rubber-stamping, no review (C) is reckless, and deep research (D) is unrelated to oversight.

  39. Q39D5 · Review Points and Human OversightSelect one

    Where should a review gate be placed in a draft-then-send workflow?

    • A. Before drafting begins
    • B. Immediately before the send, so the reviewer sees the final artefact at the last reversible moment
    • C. After the email has been sent
    • D. Placement does not matter
    Show answer

    Answer: B.

    The gate belongs just before the irreversible action, reviewing the final content. Before drafting (A) reviews nothing meaningful, after sending (C) is not a gate at all, and placement absolutely matters (D).

  40. Q40D5 · Review Points and Human OversightSelect one

    Which of these is a MANDATORY review trigger?

    • A. The output is longer than one page
    • B. The output was generated quickly
    • C. The step issues an irreversible payment
    • D. The output uses bullet points
    Show answer

    Answer: C.

    Irreversible actions like issuing a payment mandate a human gate. Length (A), speed (B) and formatting (D) are irrelevant to whether a gate is required.

  41. Q41D5 · Review Points and Human OversightSelect one

    How should the sampling rate change for a brand-new automated step?

    • A. Start low and never change it
    • B. Start with heavier sampling during burn-in, then relax as evidence of quality accumulates
    • C. Never sample new steps
    • D. Sample only after a customer complains
    Show answer

    Answer: B.

    New processes lack an evidence base, so you sample heavily at first and relax as quality is demonstrated. Starting low and never changing (A) ignores drift, not sampling new steps (C) is backwards, and waiting for complaints (D) is reactive, not oversight.

  42. Q42D5 · Review Points and Human OversightSelect one

    Why does over-gating a high-volume, low-stakes step undermine the very safety it seeks?

    • A. It makes the model hallucinate
    • B. Too many trivial approvals cause reviewer fatigue and rubber-stamping, so real errors slip through anyway
    • C. Gates always slow the model's inference
    • D. It increases token cost
    Show answer

    Answer: B.

    Excessive gates dull attention and turn review into rubber-stamping, defeating the control. It does not affect hallucination (A) or inference speed (C), and token cost (D) is not the safety issue described.

  43. Q43D5 · Review Points and Human OversightSelect two

    Which TWO steps in a workflow REQUIRE a mandatory gate rather than sampling?

    • A. Publishing a press release externally
    • B. Categorising internal receipts that reconcile monthly
    • C. Filing a regulatory disclosure
    • D. Drafting internal brainstorming notes
    • E. Summarising a meeting for your own records
    Show answer

    Answer: A and C.

    External publication and regulatory filing are irreversible/regulated/external — mandatory gates. Receipt categorisation (B) is high-volume correctable work suited to sampling, and brainstorming (D) and personal summaries (E) are low-stakes, needing at most a spot-check.

  44. Q44D5 · Review Points and Human OversightSelect one

    What makes a review gate an objective check rather than a reviewer guessing?

    • A. A bigger model
    • B. Explicit acceptance criteria from the step contract that the reviewer checks against
    • C. A longer prompt
    • D. Reviewing more items
    Show answer

    Answer: B.

    Acceptance criteria give the reviewer concrete, checkable conditions, turning review from subjective judgment into an objective check. Model size (A), prompt length (C) and volume (D) do not make a review objective.

  45. Q45D6 · Repeatability and ImprovementSelect one

    You go on leave and your backup cannot run your workflow because it only exists in your head. What is the BEST fix?

    • A. Tell them to figure it out from the chat history
    • B. Write a runbook: trigger, inputs, ordered steps with capabilities, review points, definition of done, and failure handling
    • C. Record a quick video once and hope it covers everything
    • D. Wait until you are back
    Show answer

    Answer: B.

    A written runbook is what lets a colleague run the workflow end to end without asking questions. 'Figure it out' (A) and 'wait' (D) leave the workflow un-runnable, and a one-off video (C) rarely captures inputs, failure handling and the definition of done reliably.

  46. Q46D6 · Repeatability and ImprovementSelect one

    After several undated prompt edits, a workflow got worse and no one knows which change caused it. What practice would have prevented this?

    • A. Using a bigger model
    • B. Versioning: dated versions with a one-line change note, changing one variable at a time
    • C. Writing longer prompts
    • D. Reviewing every output
    Show answer

    Answer: B.

    Dated single-change versions make effects attributable and enable rollback. A bigger model (A) and longer prompts (C) do not address traceability, and reviewing every output (D) catches errors but does not attribute which edit caused them.

  47. Q47D6 · Repeatability and ImprovementSelect one

    A manager asks whether an improved workflow is actually better. Which answer reflects evidence-based iteration?

    • A. It feels more thorough now
    • B. On a fixed 40-item sample the error rate fell from 19% to 8% and cycle time from 110 to 20 minutes
    • C. I changed a lot of things and I like the output more
    • D. The new prompt is longer
    Show answer

    Answer: B.

    Improvement is shown against measured metrics on a fixed sample. 'Feels more thorough' (A) and 'I like it more' (C) are vibes, and prompt length (D) is not a quality measure.

  48. Q48D6 · Repeatability and ImprovementSelect one

    Measuring cycle time on a workflow shows 8 of its 20 minutes are human review. What does this tell you about the next improvement?

    • A. Make the prompt longer
    • B. The biggest remaining lever is the review design, not the prompt
    • C. Switch to a cheaper model
    • D. Nothing useful
    Show answer

    Answer: B.

    Cycle-time measurement locates the bottleneck; with review dominating, review design is the lever, not prompting. Prompt length (A) and model price (C) do not touch the review time, and the measurement is highly useful (D).

  49. Q49D6 · Repeatability and ImprovementSelect two

    Which TWO are legitimate data sources for measuring a workflow's ongoing quality?

    • A. The sampled items from the review regime, scored against the acceptance criteria
    • B. The author's overall impression that it feels better
    • C. The rework rate — how often output is edited or rejected at the review gate
    • D. The number of words in the prompt
    • E. The model's release notes
    Show answer

    Answer: A and C.

    Sampled items scored against acceptance criteria, and the rework/rejection rate at the gate, are objective quality measurements. The author's impression (B) is vibes, and prompt word count (D) and release notes (E) are not quality data.

  50. Q50D6 · Repeatability and ImprovementSelect two

    Which TWO elements are essential in a runbook so a colleague can run the workflow unaided?

    • A. A definition of done
    • B. The author's personal opinion of the output
    • C. What to do when a required input is missing
    • D. The model's training cutoff date
    • E. The number of times the author has run it
    Show answer

    Answer: A and C.

    A definition of done tells the runner when to stop, and failure handling tells them what to do when input is missing — both essential for unaided execution. The author's opinion (B), the training cutoff (D) and a run count (E) do not help a colleague execute it.

Last updated Sep 18, 2026