Applied AI Foundations
Applied AI Foundations · Mock Exam 2
A second, harder 50-item independent mock exam for the Applied AI Foundations track, with new items, multi-constraint stems and a readiness indicator.
This is Mock Exam 2 for the Applied AI Foundations track: 50 new items across the six domains, none repeating Mock Exam 1 or the domain-page questions. It is an independent mock exam built from publicly available OpenAI learning objectives — not an official OpenAI assessment and not endorsed by OpenAI. Mock Exam 2 is deliberately harder: more multi-constraint stems and more FIRST / BEST / MOST cost-effective / TWO qualifiers, so you must weigh competing considerations rather than spot a single keyword. Use it as your readiness gate: reach 80%+ here, timed, before you take the real OpenAI Academy assessment.
Instructions
- Time: 60 minutes; sit it timed, under mock conditions, once you have cleared Mock Exam 1 and revised your weak domains.
- Items: 50, single-response and multiple-response. Each item states how many answers to select.
- Selection: for a select-one item choose exactly one option; for a select-two item you must choose both correct options and no incorrect one to score the item.
- No guessing penalty: answer every item — an unanswered item scores the same as a wrong one.
- Target: aim for ≥ 80% raw (≈ 40/50) before you take the real Academy assessment, which itself passes at 80%. Because this mock runs harder, a comfortable pass here is a stronger readiness signal than the same score on Mock Exam 1.
Domain distribution
| # | Domain | Weight | Items here |
|---|---|---|---|
| 1 | Finding and Scoping Opportunities | 16% | 8 |
| 2 | Decomposing Work into Steps | 18% | 9 |
| 3 | Inputs, Outputs and Contracts | 16% | 8 |
| 4 | Choosing the Right Capability | 20% | 10 |
| 5 | Review Points and Human Oversight | 16% | 8 |
| 6 | Repeatability and Improvement | 14% | 7 |
Total: 8 + 9 + 8 + 10 + 8 + 7 = 50 items.
Readiness interpretation
This is an independent readiness indicator, never an official score.
| Raw score (of 50) | Band | Interpretation |
|---|---|---|
| 40+ (80%+) | Assessment ready | On the harder mock this is a strong signal you are ready for the Academy assessment |
| 35–39 (70–79%) | Building confidence | Nearly there; target the domains where you lost the most marks |
| under 70% | Keep learning | Return to the domain pages, especially D2 and D4, before re-attempting |
| 45+ (90%+) | Strong readiness | Excellent margin on the harder items |
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it. Compare your per-domain results with Mock Exam 1 — a domain strong on Mock 1 but weak here signals shallow understanding worth revisiting.
Interactive mode
Take the practice exam
50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
Four candidates pass the value screen, but you can build only one this quarter. A saves 6 h/week with a light monthly sample; B saves 14 h/week but every output is an irreversible external filing needing per-item approval; C saves 7 h/week but touches regulated data your policy forbids leaving a sanctioned tool you do not yet have; D saves 5 h/week, fully correctable, low-sensitivity. Which should you build FIRST?
Show answer
Answer: B.
A has strong risk-adjusted payback, low risk, and no blocking dependency, so it is the best first build. B's per-item approval of irreversible filings erodes most of its 14 hours (A ignores that). C is policy-blocked until you have the sanctioned tool (C). D is buildable but its payback is the smallest of the buildable options (D).
A recurring task happens 200 times a week and saves real time, but roughly one in fifty outputs is an irreversible customer-facing action while the rest are internal and correctable. What is the MOST appropriate scoping decision?
Show answer
Answer: C.
The right move splits the stream: sample the correctable majority and gate the irreversible slice, capturing value while containing risk. Abandoning it (A) throws away the correctable value. End-to-end automation (B) puts irreversible actions in the model's hands, and a review after the fact (D) is not a gate at all.
A stakeholder insists the FIRST automation should be the CEO's quarterly investor deck because 'it is the most important thing we produce'. Which counter-argument is strongest?
Show answer
Answer: B.
Importance is not frequency: a four-times-a-year, low-tolerance showpiece is the classic impressive-but-rare trap and should be AI-assisted manually. Decks can be AI-assisted (A). Length (C) and a presumed objection (D) are not the analytical reason it fails the screen.
You have scoped a workflow to 'draft first-reply text for known ticket types'. Over three weeks it quietly grows to also tag priority, send the reply, and issue small credits. What scoping element failed, and what is the fix?
Show answer
Answer: B.
Silent expansion into sending and issuing credits is scope creep that a written out-of-scope boundary is meant to prevent. The trigger (A) governs when it runs, not what it must avoid. Model size (C) is irrelevant, and unmanaged creep into irreversible actions (D) is a risk, not healthy growth.
Which TWO characteristics most strongly signal that a task should keep a human fully in the loop rather than be automated end-to-end, even when it is frequent and time-consuming?
Show answer
Answer: A and C.
Irreversibility and regulated legal exposure are the risk signals that keep a human in the loop regardless of volume. Internal correctable output (B) and public data (D) lower risk, and tedium (E) is a reason to automate, not to keep it manual.
Two tasks each save about 10 hours a week. Task P is fully repeatable end to end; task Q has a large, variable human-judgment core that must be exercised on every instance. Which is the BETTER automation candidate and why?
Show answer
Answer: B.
The realizable saving depends on how much of the task is a stable repeatable core; Q's large variable judgment part cannot be automated, so less of its 10 hours is actually removed. Judgment tasks are not automatically more valuable (A), the two are not identical once you separate core from judgment (C), and 10 hours is a substantial saving (D).
A one-time migration will take three days of tedious manual reformatting. A colleague proposes building a reusable workflow for it. What is the MOST cost-effective decision?
Show answer
Answer: B.
Frequency of one means a productised workflow never runs again, so ad hoc AI assistance captures the saving without the build overhead. Building a reusable workflow (A) or scheduling it monthly (D) wastes effort on a task that will not recur, and doing it by hand (C) forgoes the available AI help.
A teammate proposes fixing an unreliable mega-prompt that produces a launch pack by (1) upgrading to the most expensive model AND (2) adding three more paragraphs of instructions. Why is decomposition the BETTER first move?
Show answer
Answer: B.
The problem is structural, not capability; decomposing exposes and isolates failures and often lets lighter models run each step. The expensive model is not banned (A), cheaper models are not always more accurate (C), and decomposition does not remove review (D).
You are designing a national brief from twelve regional reports that must reconcile terminology and totals, and each report buries the figures in prose. Which composed structure is BEST?
Show answer
Answer: B.
Independent reports fan out (each extracted to a mini-structure) and the fan-in synthesises with a critique that reconciles terms and totals — patterns composed to fit. Pure sequential (A) ignores independence, critiquing all twelve raw (C) skips per-report structure, and one mega-prompt (D) is the anti-pattern.
A workflow was split into a dozen tiny steps, each doing a fraction of a job, and it is now slow and loses context between handoffs while adding no new checks. What is the BEST corrective action?
Show answer
Answer: B.
Over-splitting adds latency and context loss without new inspection, so the fix is to merge steps that share one job while retaining meaningful splits. More steps (A) worsens it, one mega-prompt (C) is the opposite anti-pattern, and a bigger model (D) does not remove handoffs.
A contract-review workflow drafts a memo then checks it in the same reply, and reviewers notice the self-check almost never flags anything. What single change MOST improves the critique's objectivity?
Show answer
Answer: B.
A same-pass self-review defends its own draft; a separate checklist-driven critique step is the fix that makes review objective. 'Try harder' (A) is not a structural change, token limits (C) do not affect objectivity, and drafting temperature (D) does not fix the missing separate pass.
Which ordering is correct for a robust client proposal built from a discovery transcript, where numbers must be checked and the summary must match the final body?
Show answer
Answer: B.
Extract first, produce the numbers with a check where they are created, draft from those, summarise last, then check against terms — each step inspectable and the summary matching the final body. Summary-first (A) drifts, one prompt (C) is the mega-prompt, and pricing before scope (D) has no inputs to price.
A task involves reading a dense supplier email and then deciding, from the extracted terms, whether it breaches a policy threshold. A single prompt sometimes flags a breach the numbers do not support. What is the ROOT-CAUSE decomposition?
Show answer
Answer: B.
Blending extraction and the threshold decision hides whether an error is a misread term or a misapplied rule; extract-then-transform separates them for checking. Draft-then-critique (A) reviews prose, a longer single answer (C) keeps the errors blended, and fanning over paragraphs (D) does not separate extraction from the rule.
A single mega-prompt produces a five-part deliverable that drops parts and contradicts itself. Which TWO structural fixes address these two distinct symptoms at the root?
Show answer
Answer: A and C.
Named fan-out steps make dropped parts visible, and a shared extracted fact sheet removes the contradictions at the root. An instruction not to drop parts (B) is prompt padding, a bigger model (D) hides rather than exposes failures, and generating twice (E) doubles cost without fixing either cause.
A downstream billing step requires a numeric total, but the producing step's contract only promises a formatted string like '$1,204.00'. What is the BEST fix?
Show answer
Answer: B.
The clean fix aligns the producer's output type to the consumer's required numeric input rather than relying on fragile downstream parsing. Parsing and hoping (A) is brittle to format changes, a bigger model (C) does not fix a type mismatch, and 'be more consistent' (D) is not a contract change.
An extraction step outputs sentiment as free text like 'mostly positive' instead of the expected enum, and a colleague running it gets yet other phrasings. Which change MOST directly fixes both problems?
Show answer
Answer: B.
An enumerated value set in the schema forces one of the allowed values for every user and input; free text means the value set was never constrained. 'Be careful' (A) is not a contract, max tokens (C) is unrelated, and more examples (D) do not constrain the output enum.
A step that classifies feedback outputs a product area even when the email names no product, polluting the CRM with guessed values. What is the BEST contract change?
Show answer
Answer: B.
An explicit missing-input rule (unknown plus a review flag) combined with a constrained value set stops the fabrication. Guess-and-log (A) still pollutes with fabricated values, removing the field (C) drops a required output, and higher temperature (D) worsens variance.
You are choosing the output shape for a step whose result a human reads and must check for completeness. Which shape is MOST appropriate?
Show answer
Answer: B.
For a human-read deliverable checked for completeness, a fixed template makes a missing section visible. A strict JSON schema (A) suits a parser, not a reader, and free-form prose (C) or a single paragraph (D) makes completeness hard to verify.
A colleague says 'the output looked fine in my three tests, so the shape is defined'. Why is this reasoning unsafe?
Show answer
Answer: B.
A handful of passing samples does not pin the shape; the format can drift on unseen inputs or with another user unless the contract specifies it. There is no such policy (A), no retraining requirement (C), and testing does not prove a contract is complete (D).
A required 'known_customers' reference is absent when a lookup step runs, and correctness fully depends on it. What is the correct missing-input behaviour?
Show answer
Answer: B.
When correctness depends on a value and no safe default exists, the step must fail loudly rather than fabricate. A most-common default (A), a guess (C), and an alphabetical pick (D) all substitute invented data for a required input.
Why does writing the step contract matter MOST for handover and repeatability?
Show answer
Answer: B.
A written contract is the runnable interface a colleague follows to get the same result; an undocumented step lives only in the author's head. It does not shorten prompts (A), skip criteria (C), or remove review (D).
You are hardening a two-step chain where step 1 feeds step 2. Which TWO checks MOST reduce flakiness at the boundary?
Show answer
Answer: A and C.
Matching the producer's output to the consumer's required input, and defining step 2's missing/malformed-input behaviour, are the two boundary checks that remove most flakiness. Matching temperature (B), prompt length (D) and run day (E) are not contract-boundary concerns.
A manager says 'just build one API application' for a bundle that is (a) your own recurring digest, (b) a version reps run themselves in ChatGPT, and (c) a future scheduled Slack post. What is the BEST right-sizing?
Show answer
Answer: B.
Each need maps to a different rung: recurring solo context is a Project, others self-serving is a custom GPT, and scheduled integrated automation is an API application. One API app for all (A) over-builds (a) and (b), a custom GPT for all (C) cannot do scheduled Slack posting, and weekly prompts (D) neither scale nor self-serve.
A stem says the output 'must be a sourced, cited market report that is THEN refined into a polished brief over several edits'. Which single choice best matches the FULL requirement?
Show answer
Answer: B.
A sourced, cited report calls for deep research, and refining it across several edits calls for Canvas — both parts of the requirement. Search (A) under-delivers the report, data analysis (C) fits no file-computation need here, and a one-question agent (D) does not match a multi-edit deliverable.
You keep re-pasting a 4-page background into fresh chats for a monthly task and only you run it. Which is the MOST cost-effective fix?
Show answer
Answer: B.
Recurring standing context for a solo task is exactly a Project, the lightest rung that removes the re-pasting. An API app (A) and a custom GPT (C) over-reach when only you run it, and a bigger model (D) does not fix the re-pasting or the missing standing context.
A workflow currently done as weekly prompts must eventually run automatically every Monday and post to a Slack channel. Which rung does THAT specific requirement justify, and why?
Show answer
Answer: C.
Scheduled, automated, integrated-with-Slack execution is embedding at scale, the justification for an API application. A saved instruction (A) and a Project (B) do not run on a schedule wired into Slack, and a custom GPT (D) is for others to self-serve, not scheduled posting.
A colleague reaches for deep research to answer 'what is our competitor's stock price right now?'. What is the BEST correction?
Show answer
Answer: B.
A single current fact is a search job; deep research is heavier and reserved for cited, multi-source synthesis. Deep research here (A) wastes time, data analysis (C) needs a file to compute over, and Canvas (D) is an editing surface, not a lookup tool.
Two tasks are proposed: (1) a repeatable assistant that colleagues chat with turn by turn, and (2) a multi-step task delegated to run semi-autonomously with checkpoints. Which mapping is correct?
Show answer
Answer: B.
A turn-by-turn assistant others chat with is a custom GPT, while delegated multi-step execution under oversight is a workspace agent. Swapping them (A) inverts answer-vs-execute, and neither is inherently an API app (C) or a one-off prompt (D).
A Project that produces a weekly digest has started giving stale, subtly wrong outputs even though the prompt is unchanged. What is the MOST likely cause?
Show answer
Answer: B.
Heavier rungs carry a maintenance cost; stale Project knowledge files silently mislead the output even when the prompt is untouched. Models do not silently downgrade (A), temperature does not drift on its own (C), and Projects are perfectly capable of reliable output when maintained (D).
Which TWO tasks are best served by an in-conversation capability rather than by moving to a heavier packaging rung?
Show answer
Answer: A and C.
Computation over a file (data analysis) and a current-facts lookup (search) are in-conversation capabilities. Fifty colleagues self-serving (B) is a custom GPT, and nightly scheduling (D) or embedding in a product UI (E) are API-application needs.
A team is about to build a custom GPT for a task that, on inspection, only its author ever runs and simply needs the same three files each time. What is the BEST advice?
Show answer
Answer: B.
When only the author runs it and it needs standing context, a Project is the correct lighter rung; a custom GPT is justified only when others run it. Proceeding with the GPT (A) over-reaches, an API app (C) over-reaches further, and re-pasting (D) is the waste a Project removes.
A support workflow auto-sends replies for 'simple' categories. One category quotes a delivery date pulled from a field that can be stale, and a wrong date reached a customer. What is the BEST targeted fix?
Show answer
Answer: B.
The risk is specific to replies depending on stale-able data, so gate or verify those while leaving genuinely self-contained replies automated. Gating everything (A) is the throttling over-correction, abandoning AI (C) is disproportionate, and temperature (D) does not fix stale data.
A workflow shows 97% aggregate accuracy, so a manager proposes removing the gate before an irreversible legal filing. What is the BEST response?
Show answer
Answer: B.
A high average does not cover the low-frequency, high-consequence errors on an irreversible regulated step, so the mandatory gate stays. 97% 'good enough' (A) ignores the tail, a gate after filing (C) is no gate, and a higher target (D) still cannot make an irreversible regulated action safe to leave ungated.
A reviewer is approving 500 items an hour and only glancing at formatting. Which TWO changes MOST improve oversight without simply doing more work?
Show answer
Answer: A and B.
Fewer, meaningful gates via sampling, plus criteria-driven substantive review, fix rubber-stamping. Doubling the load (C) worsens fatigue, removing review (D) is reckless, and a bigger model (E) does not eliminate the need for oversight on risky steps.
A gate has been placed on the draft produced early in a workflow, but the draft is substantially rewritten by two later steps before the irreversible send. What is wrong with this placement?
Show answer
Answer: B.
A gate must review the final artefact at the last reversible moment; reviewing an early draft that later changes wastes the review. Earlier is not always better (A), the gate is needed (C), and the draft step is legitimate (D) — only the gate's placement is wrong.
For a high-volume categorisation step, sampled error rate has been climbing for two weeks on high-value items specifically. What is the BEST response?
Show answer
Answer: B.
Drift concentrated in a risky slice calls for tighter sampling and a targeted gate on that slice, not blanket action. Doing nothing (A) ignores the drift, stopping sampling (C) removes the signal, and gating every item (D) is the throttling over-correction.
Which mode-to-step mapping is correct?
Show answer
Answer: B.
Grave steps are gated, high-volume correctable work is sampled, and stable low-stakes work is spot-checked. The other options invert the mapping, e.g. sampling an irreversible publication (A), spot-checking an irreversible payment (C), or sampling a regulatory filing (D).
A team over-reacted to one bad auto-sent reply by gating all 20,000 daily replies. Which TWO consequences make this over-correction self-defeating?
Show answer
Answer: A and C.
A blanket gate on high-volume work throttles throughput and induces rubber-stamping, defeating the control. It does not change model accuracy (B), violate policy (D), or self-document the workflow (E).
A false-positive regression appeared after many untracked prompt edits. You now adopt versioning. What is the BEST recovery?
Show answer
Answer: B.
Rolling back to a known-good version and re-applying changes one at a time with measurement isolates the offending edit and restores quality. Rebuilding from memory (A) loses the good history, keeping the edits and changing temperature (C) does not isolate the cause, and adding rules (D) compounds the noise.
A support-reply prompt was edited to 'be more concise' and customer satisfaction dropped because replies now omit a needed next step. With versioning in place, what is the BEST response?
Show answer
Answer: B.
Versioning lets you revert the regression and re-apply a single isolated change. Starting over (A) discards working history, keeping a version that measurably hurt satisfaction (C) is not evidence-based, and knowledge files (D) do not address the omitted next step.
Before optimising a workflow, you measure that 12 of its 18 minutes per run are spent in a manual review step. What is the MOST effective next move?
Show answer
Answer: B.
Cycle-time measurement locates the bottleneck; with review consuming most of the run, redesigning review is the biggest lever. Shortening the prompt (A) or adding automated steps (D) does not touch the dominant review time, and model price (C) is not the bottleneck here.
Which TWO properties distinguish evidence-based iteration from vibes-based 'improvement'?
Show answer
Answer: A and B.
Single-variable changes plus metric-based keep-or-revert with rollback are what make iteration attributable and reversible. Changing several settings at once (C) and judging by feel (D) are the vibes anti-pattern, and evidence-based iteration keeps changes that improve the metric (E).
A workflow cut a task from 120 to 20 minutes per run. What does this measurement PRIMARILY support?
Show answer
Answer: B.
Cycle-time savings quantify the payback screened at the opportunity stage and give a baseline to improve against. Time is highly relevant (A), it does not by itself prove quality (C) — that needs the acceptance-criteria metric — and it has nothing to do with temperature (D).
A colleague offers 'record a video of me doing it once' as the documentation for handover. Why is a runbook the BETTER choice?
Show answer
Answer: B.
A runbook is the structured, executable record of the workflow, whereas a one-off video usually misses inputs, failure handling and the definition of done. There is no such policy (A), videos can be stored (C), and a runbook records review points rather than removing them (D).
Which of these is the correct order of the DRIVE maintenance loop for a live workflow?
Show answer
Answer: B.
DRIVE is Document, Record versions, Instrument, Vary one thing, Evaluate — a disciplined single-change, measured cycle. The other options either change many variables at once, judge by feel, or skip measurement, which is the vibes anti-pattern.
A weekly workflow summarises ten separate customer-call transcripts and must produce one consistent themes report. A colleague simply concatenates the ten summaries. What is the MOST accurate critique?
Show answer
Answer: B.
Fan-in is a synthesis step that resolves contradictions, dedupes themes and reconciles across branches; concatenation skips the actual combining work. Concatenation is not fan-in (A), the transcript count (C) is irrelevant, and running the independent transcripts sequentially (D) would misfit the parallel structure.
A single prompt is asked to (i) extract figures from an invoice and (ii) apply reimbursement rules, and wrong totals cannot be traced to either cause. Which TWO changes best restore traceability?
Show answer
Answer: A and C.
Splitting into an extraction step and a rule-application step (extract-then-transform) makes each independently checkable, so a wrong total traces to reading or logic. A self-check sentence (B) keeps the errors blended, and a bigger model (D) or more tokens (E) does not separate the two failure sources.
A task needs the same reference files each run AND must compute totals over an uploaded spreadsheet each time. What is the BEST combination?
Show answer
Answer: B.
The packaging need (standing files, solo) is a Project, and the compute need is the in-conversation data analysis capability — the two axes combine. A one-off prompt (A) loses the standing files, a custom GPT (C) over-reaches when only you run it, and an API app (D) is unjustified merely because a spreadsheet is involved.
A workflow drafts and then a human sends replies, but leadership wants to add a second gate 'just in case' immediately after the send. Why is this the WRONG design?
Show answer
Answer: B.
Review after an irreversible action is not a gate, because the harm has already occurred; the effective gate is the one before the send. More gates are not always better (A), the pre-send gate should stay (C), and post-send review does not affect model accuracy (D).
A candidate task scores high on frequency and time cost, but its output must be published externally with almost no error tolerance. On the FTED screen, what is the MOST appropriate build decision?
Show answer
Answer: B.
High value with low error tolerance means you capture the value by automating the drafting but keep a mandatory gate before the external, hard-to-retract action. Skipping (A) throws away real value, full automation (C) puts an unforgiving external step in the model's hands, and model price (D) does not address the error-tolerance risk.
Last updated Sep 18, 2026