AI Cert Prep
Type to search documentation.

AI Foundations

AI Foundations Mock Exam 2

A harder, timed 60-item independent mock exam for the OpenAI Academy AI Foundations track, used as a readiness gate before the real assessment.

This is the second full-length independent mock exam for the AI Foundations track, built from publicly available OpenAI learning objectives. It is not an official OpenAI assessment. It is deliberately harder than Mock Exam 1 – more multi-constraint stems, more FIRST / BEST / MOST cost-effective / TWO qualifiers and more scenario framing – so use it as your readiness gate, timed, once you have closed the gaps the diagnostic exposed. All 60 items are new and distinct from Mock Exam 1 and the domain-page questions.

Instructions

  • Time: 75 minutes. Take this one timed to rehearse working under a clock.
  • Items: 60, single-answer and multiple-response. Each item states how many answers to select.
  • Selection: single-answer items use one choice; multiple-response items say Select two and you must pick exactly two – all correct and none incorrect.
  • No guessing penalty: answer every question; an unanswered item simply scores zero.
  • Target: aim for at least 80% raw (≈ 48/60) here before you take the real Academy assessment, because the Academy badge threshold is 80%.

Domain distribution

#DomainItems here
1AI and Generative AI Fundamentals8
2How Language Models Behave10
3ChatGPT Surfaces and Features8
4Prompting and Instructions12
5Context, Files and Memory7
6Verifying and Evaluating Output9
7Responsible and Safe Use6

Total: 8 + 10 + 8 + 12 + 7 + 9 + 6 = 60 items.

Readiness interpretation

This is an independent readiness indicator, not a score and not a pass mark. Because this exam is harder, a score in the upper bands here is a strong signal of readiness.

Raw scoreBandWhat it means
under 70%Keep learningThe harder framing found real gaps; revisit those domains before the assessment.
70–79%Building confidenceSolid, but tighten the domains where multi-constraint items caught you.
80–89%Assessment readyAt or above the Academy badge threshold on the harder set; you are in good shape.
90%+Strong readinessExcellent on the tougher exam; you are well prepared for the Academy assessment.

The 80% line matches the OpenAI Academy badge threshold. An Academy badge or pathway certificate of completion is not a certification and does not guarantee eligibility for a future OpenAI certification.

Take the mock exam

Two ways to use the questions below. The interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction. The review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

60 questions · one at a time · 75-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · AI and Generative AI FundamentalsSelect one

    A stakeholder claims that because ChatGPT wrote correct Python, it must therefore understand mathematics the way a person does. What is the MOST accurate correction?

    • A. Correct; producing working code proves human-like understanding.
    • B. The model produced statistically probable, plausible-looking text that happened to be valid; producing fluent output is not the same as human comprehension.
    • C. It understands mathematics but not language.
    • D. It memorised the exact program from training and replayed it verbatim.
    Show answer

    Answer: B.

    The model generates probable tokens; correct output demonstrates capability, not human-style understanding. A overclaims. C is backwards, since it is a language model. D assumes rote replay, which mischaracterises generative next-token prediction and is usually untrue for novel prompts.

  2. Q2D1 · AI and Generative AI FundamentalsSelect one

    You must estimate whether a 6,000-word briefing plus a 2,000-word appendix will fit alongside instructions in a GPT Instant window on ChatGPT Business. Using the standard rule of thumb, what is the BEST estimate and conclusion?

    • A. About 2,700 tokens, comfortably within the 54K Business Instant window.
    • B. About 10,700 tokens, comfortably within the 54K Business Instant window with room for instructions.
    • C. About 10,700 tokens, which exceeds the 54K Business Instant window.
    • D. About 32,000 tokens, exceeding the window so it cannot be used at all.
    Show answer

    Answer: B.

    At roughly four characters per token, 8,000 words is about 10,700 tokens, well under 54K, leaving room for instructions. A underestimates by using far too few tokens per word. C wrongly claims 10,700 exceeds 54K. D both overestimates the count and misstates the limit.

  3. Q3D1 · AI and Generative AI FundamentalsSelect one

    A team asks for the SINGLE most reliable general explanation of why an LLM sometimes miscounts letters in a word. Which is it?

    • A. It works over tokens, not individual characters, so character-level counting is not native to how it processes text.
    • B. It is deliberately programmed to miscount as a safety feature.
    • C. Its knowledge cutoff removed spelling data.
    • D. The context window is always too small for words.
    Show answer

    Answer: A.

    Models operate on tokens rather than characters, so exact letter counting is not a natural operation. B invents a feature. C confuses training recency with character handling. D misapplies context limits to a per-word counting task.

  4. Q4D1 · AI and Generative AI FundamentalsSelect one

    A knowledge worker wants the MOST cost-effective way to handle a repeatable, high-volume extraction job with clear rules, and asks which factor matters FIRST. What should drive the choice?

    • A. Always pick the most capable frontier model for safety.
    • B. Match a lightweight, fast model to the clear, repeatable task and reserve heavier models for genuinely hard work.
    • C. Pick a realtime voice model for throughput.
    • D. Choose based on which model has the newest knowledge cutoff.
    Show answer

    Answer: B.

    For clear, repeatable, high-volume tasks, a lightweight fast model is the cost-effective fit. A overspends. C uses the wrong modality. D optimises for recency, which is irrelevant to a rules-based extraction task.

  5. Q5D1 · AI and Generative AI FundamentalsSelect one

    Which statement BEST captures why fluency is not the same as accuracy in generated text?

    • A. Fluency is a measure of factual correctness computed during generation.
    • B. The model optimises for probable, well-formed wording, which is independent of whether the underlying claims are true.
    • C. Only slow models can be inaccurate.
    • D. Accuracy is guaranteed whenever a citation is present.
    Show answer

    Answer: B.

    Well-formed wording reflects probability, not truth, so fluent text can be false. A wrongly equates fluency with correctness. C ties accuracy to speed. D is false, since citations themselves can be fabricated.

  6. Q6D1 · AI and Generative AI FundamentalsSelect two

    A colleague asks you to name TWO facts that BEST explain why a precise, unsourced statistic about a small private company should be treated with caution.

    • A. The model can generate a specific-looking number that it did not retrieve from any source.
    • B. Niche private data is often outside training data and unavailable without a search tool.
    • C. Precise numbers are always verified by the model before output.
    • D. The model always cites a source for statistics.
    • E. Small companies are never mentioned in any text.
    Show answer

    Answer: A and B.

    The model can fabricate precise figures, and niche private data is often unavailable without search. C is false; there is no built-in verification. D is false; unsourced numbers are common. E is an overstatement that is not why caution is warranted.

  7. Q7D1 · AI and Generative AI FundamentalsSelect one

    A manager argues that adopting AI means the company should stop hiring analysts because the model computes everything reliably. What is the MOST accurate framing to offer FIRST?

    • A. Agreed; the model is a reliable calculator and analyst replacement.
    • B. The model is a fast, fallible collaborator that generates probable text; consequential analysis still needs human judgment and tools for exact computation.
    • C. The model cannot help with analysis at all.
    • D. It only fails on tasks after its knowledge cutoff.
    Show answer

    Answer: B.

    The model assists but does not replace human judgment, and exact computation needs tools. A overclaims reliability. C understates its usefulness. D wrongly limits failure modes to recency, ignoring hallucination and arithmetic errors.

  8. Q8D1 · AI and Generative AI FundamentalsSelect one

    For a task that must reflect a regulatory change made three months ago, which combination is the MOST dependable FIRST choice?

    • A. Rely on the base model's training knowledge alone.
    • B. Use a search-enabled feature to retrieve the current rule, then verify against the primary source.
    • C. Raise the reasoning effort to the maximum.
    • D. Ask the model to guess the most likely rule.
    Show answer

    Answer: B.

    Recent facts need retrieval plus verification against the authoritative source. A risks stale or fabricated content past the cutoff. C adds compute but cannot supply missing recent facts. D invites a confident guess with no grounding.

  9. Q9D2 · How Language Models BehaveSelect one

    A user sets what they think is a temperature of zero and declares the model can no longer hallucinate or vary. What is the MOST accurate correction?

    • A. Correct; zero temperature removes hallucination and variation entirely.
    • B. Lower randomness can reduce wording variation but does not eliminate hallucination, and ChatGPT does not expose a raw temperature control to end users.
    • C. Temperature only affects speed.
    • D. Zero temperature switches the model to a database lookup.
    Show answer

    Answer: B.

    Reducing randomness affects variation, not factual grounding, and the chat product does not surface raw temperature. A conflates determinism with truth. C misstates what the setting does. D invents a lookup mode that does not exist.

  10. Q10D2 · How Language Models BehaveSelect one

    A user wants to reliably retrieve an excellent draft they generated last week, but the chat is gone. What is the BEST practice going forward?

    • A. Rely on the model regenerating the identical text on request.
    • B. Save valuable outputs deliberately, because non-deterministic generation will not reproduce them exactly.
    • C. Increase reasoning effort so outputs become reproducible.
    • D. Assume memory stored the full draft automatically.
    Show answer

    Answer: B.

    Because generation is non-deterministic, valuable outputs must be saved intentionally. A expects exact reproduction that will not happen. C does not make outputs reproducible. D wrongly assumes memory stores full drafts, which it does not.

  11. Q11D2 · How Language Models BehaveSelect one

    On ChatGPT Enterprise, a user compares the GPT Instant and GPT Reasoning total context windows and asks which pairing is correct.

    • A. 128K for Instant and 256K for Reasoning.
    • B. 54K for Instant and 256K for Reasoning.
    • C. 256K for Instant and 128K for Reasoning.
    • D. 1M for both.
    Show answer

    Answer: A.

    On Enterprise the GPT Instant total is 128K and Reasoning is 256K. B uses the Business Instant figure. C swaps the two values. D quotes an API-scale window not applicable to these ChatGPT plans.

  12. Q12D2 · How Language Models BehaveSelect one

    A user needs BEST quality on a hard, open-ended strategy problem but is cost-conscious. Which selection reasoning is soundest?

    • A. Always use the cheapest model regardless of difficulty.
    • B. Use a capable reasoning-oriented model at a reasoning effort matched to the difficulty, not the maximum by reflex.
    • C. Use the maximum effort on the smallest model.
    • D. Use a voice model for strategy.
    Show answer

    Answer: B.

    Hard open-ended work needs a capable model, but effort should match difficulty rather than defaulting to maximum. A under-provisions for a hard task. C pairs high effort with an underpowered model. D uses the wrong modality entirely.

  13. Q13D2 · How Language Models BehaveSelect one

    A model gives a fast, wrong answer to a chained arithmetic-and-logic problem. A user increases the reasoning effort and it improves but is still occasionally wrong. What is the MOST reliable next step?

    • A. Accept the improved answer as good enough without checking.
    • B. Have the numeric parts computed by a data-analysis tool and verify the logic, rather than relying on generation alone.
    • C. Lower the reasoning effort to speed it up.
    • D. Re-send the prompt many times and take a majority vote.
    Show answer

    Answer: B.

    Exact computation belongs in a tool, with the logic verified, since more effort alone does not guarantee arithmetic correctness. A skips verification on a task known to fail. C reduces rigour. D relies on repeated fallible generations rather than grounded computation.

  14. Q14D2 · How Language Models BehaveSelect one

    A user notices the model enthusiastically endorses whichever option they frame favourably. To get the MOST balanced decision support, what should they do FIRST?

    • A. Ask leading questions to confirm the preferred option.
    • B. Ask for a neutral comparison with explicit pros and cons for every option and the strongest case against their favourite.
    • C. Increase the context window.
    • D. Switch to image generation.
    Show answer

    Answer: B.

    Neutral framing that forces balanced pros, cons and counterarguments counters sycophancy. A amplifies the bias. C adds capacity irrelevant to framing. D changes modality and does not address the reasoning need.

  15. Q15D2 · How Language Models BehaveSelect one

    In a marathon planning chat, the model starts contradicting decisions made hours earlier. Which explanation and remedy is MOST accurate?

    • A. The model is broken; report it and wait.
    • B. Early turns have aged out of the effective context; summarise the key decisions and carry them forward or start a fresh, focused chat.
    • C. The knowledge cutoff advanced during the chat.
    • D. The plan exceeded the monthly message quota.
    Show answer

    Answer: B.

    Long chats lose early context; carrying forward a summary or restarting preserves the decisions. A misreads normal behaviour. C confuses training cutoff with in-chat context. D is a billing concept unrelated to contradictions.

  16. Q16D2 · How Language Models BehaveSelect two

    Select the TWO statements that correctly describe consequences of probabilistic, non-deterministic generation for everyday ChatGPT use.

    • A. Identical prompts can yield differently worded answers across runs.
    • B. A valuable output should be saved because it may not be reproduced exactly.
    • C. Setting a preference guarantees byte-identical outputs forever.
    • D. The model verifies every fact before generating it.
    • E. Non-determinism prevents the model from ever hallucinating.
    Show answer

    Answer: A and B.

    Non-determinism explains varied wording and the need to save good outputs. C is false; preferences do not fix exact bytes. D is false; there is no built-in fact verification. E is false; variability and hallucination are separate phenomena.

  17. Q17D2 · How Language Models BehaveSelect one

    A user reports that quality degraded after they pasted a dozen long, loosely related reports into one chat, though a token counter shows they are under the window. What is the MOST likely cause and fix?

    • A. They exceeded the window; delete half the reports at random.
    • B. Relevant signal is diluted by irrelevant material; attach only what the task needs, ideally organised in a Project.
    • C. The model cannot process reports; convert them to images.
    • D. Raise reasoning effort to force it to read everything.
    Show answer

    Answer: B.

    Under the limit, relevance and volume still matter; pruning to relevant files fixes dilution. A contradicts the premise and prunes blindly. C invents a limitation. D adds compute without reducing the clutter causing the problem.

  18. Q18D3 · ChatGPT Surfaces and FeaturesSelect one

    A user must confirm a competitor price announced this morning, then produce an exact side-by-side cost table from a 300-row sheet. Which pairing of features is correct, in order?

    • A. Memory to recall the price, then image generation for the table.
    • B. Search to confirm the fresh price, then data analysis to compute the table exactly.
    • C. Deep research for the price, then canvas to calculate totals.
    • D. Study mode for both steps.
    Show answer

    Answer: B.

    Search handles the fresh fact and data analysis computes exact figures. A misuses memory and image generation. C uses deep research for a one-line fact and wrongly expects canvas to compute. D is a tutoring feature that does neither task.

  19. Q19D3 · ChatGPT Surfaces and FeaturesSelect one

    A team wants a shared, standing configuration of instructions and reference files that every member's chats can draw on for one client, without pasting anything. Which is the BEST fit?

    • A. A memory entry per person.
    • B. A shared Project with custom instructions and uploaded files.
    • C. A scheduled task.
    • D. A single long chat everyone appends to.
    Show answer

    Answer: B.

    A shared Project holds standing instructions and files across members' chats. A is per-person and small-scale. C automates timing, not shared context. D degrades as it grows and does not cleanly share configuration.

  20. Q20D3 · ChatGPT Surfaces and FeaturesSelect one

    A user needs a well-sourced, multi-article synthesis on a niche standard AND wants each claim traceable to a source. Which feature and follow-up is BEST?

    • A. A single fast reply, trusting its unsourced claims.
    • B. Deep research to gather cited sources, followed by manually opening and confirming the key citations.
    • C. Image generation with captions.
    • D. Voice mode dictation.
    Show answer

    Answer: B.

    Deep research synthesises cited material, and confirming key citations makes claims traceable. A offers no sourcing. C and D are the wrong modalities and provide no traceable synthesis.

  21. Q21D3 · ChatGPT Surfaces and FeaturesSelect one

    A manager insists every task run on the most powerful model to be safe. Which single point BEST pushes back on cost-effectiveness grounds?

    • A. The most powerful model never costs more.
    • B. For clear, repeatable tasks a lighter model gives equivalent results faster and cheaper, so effort and model should match the task.
    • C. Powerful models cannot do simple tasks.
    • D. Model choice has no effect on cost or speed.
    Show answer

    Answer: B.

    Matching model and effort to task avoids paying frontier prices for simple work. A is false on cost. C is false; capable models handle simple tasks but wastefully. D denies the real cost and latency differences.

  22. Q22D3 · ChatGPT Surfaces and FeaturesSelect one

    An analyst wants to learn a forecasting method deeply rather than be handed the final numbers, then later needs the numbers computed exactly. Which sequence is BEST?

    • A. Study mode to learn the method, then data analysis to compute the numbers.
    • B. Search for the method, then memory to store the numbers.
    • C. Image generation for both.
    • D. One fast reply asking for the answer directly.
    Show answer

    Answer: A.

    Study mode teaches the method; data analysis computes exact numbers with code. B uses search for a concept and misuses memory for computation. C is the wrong modality. D skips learning and risks unreliable arithmetic.

  23. Q23D3 · ChatGPT Surfaces and FeaturesSelect two

    Select the TWO correctly matched feature-to-need pairs.

    • A. Answering from a connected internal document store uses company knowledge.
    • B. Iteratively editing a single long document inline uses canvas.
    • C. Guaranteeing exact totals of a 500-row sheet uses memory.
    • D. Confirming this morning's news uses image generation.
    • E. Storing a 50-page PDF for a one-time summary uses a scheduled task.
    Show answer

    Answer: A and B.

    Company knowledge answers from connected internal sources and canvas supports inline editing. C should use data analysis, not memory. D should use search, not image generation. E should use a file upload, not a scheduled task.

  24. Q24D3 · ChatGPT Surfaces and FeaturesSelect one

    A user uploads a 60-page PDF and asks for a summary that states a specific compliance figure they will forward to legal. What is the correct posture BEFORE forwarding?

    • A. Forward it; the model read the document, so it must be right.
    • B. Locate the figure in the PDF and confirm the summary quotes it correctly before forwarding.
    • C. Ask the model whether it is confident.
    • D. Round the figure before sending.
    Show answer

    Answer: B.

    A high-stakes figure destined for legal must be checked against the source document. A over-trusts summarisation. C relies on self-reported confidence. D alters the figure without confirming it.

  25. Q25D4 · Prompting and InstructionsSelect one

    A task prompt says: you are a financial analyst; using the attached CSV, report the three largest cost categories as a table with columns Category and Total; success means the totals reconcile to the sheet. Which components are present?

    • A. Role, task, context (the CSV), format and a success criterion.
    • B. Only a role and a task.
    • C. Only a format, with no role or task.
    • D. Only context and audience.
    Show answer

    Answer: A.

    The prompt names a role, a task, the CSV as context, a table format and a reconciliation success criterion. B, C and D each undercount the components clearly present in the stem, so they misidentify the well-formed prompt.

  26. Q26D4 · Prompting and InstructionsSelect one

    An output must feed an automated importer that expects fixed fields. The first attempt returned prose with the data embedded. What is the BEST single instruction to add?

    • A. Ask it to be more concise.
    • B. Specify the exact structured schema, for example JSON with named fields, and require no prose outside it.
    • C. Raise the reasoning effort.
    • D. Request a friendlier tone.
    Show answer

    Answer: B.

    A defined schema with no extra prose makes output importable. A shortens prose but keeps it unstructured. C adds compute without defining structure. D changes tone, which is irrelevant to machine import.

  27. Q27D4 · Prompting and InstructionsSelect one

    A prompt lists five prohibitions and no positive target; outputs keep missing the mark. Applying prompt-anatomy thinking, what is the MOST effective single fix?

    • A. Add a sixth prohibition.
    • B. Replace the prohibitions with a positive description of the desired output plus one worked example.
    • C. Shorten the prompt to one line.
    • D. Increase temperature for creativity.
    Show answer

    Answer: B.

    A positive target with an example steers far better than stacked prohibitions. A adds more negatives without a target. C removes needed detail. D adds randomness, not direction.

  28. Q28D4 · Prompting and InstructionsSelect one

    A complex deliverable needs research, a recommendation, a slide outline and a summary email, each depending on the previous. What is the MOST reliable prompting structure?

    • A. One mega-prompt that asks for all four artefacts together.
    • B. A sequence of prompts, each producing and reviewing one artefact before feeding it into the next.
    • C. Ask only for the email and infer the rest.
    • D. Randomly reorder the steps to test robustness.
    Show answer

    Answer: B.

    Sequential, reviewed steps let errors be caught before they propagate through dependent artefacts. A makes review and error isolation hard. C skips required work. D introduces disorder without benefit.

  29. Q29D4 · Prompting and InstructionsSelect one

    A user reports that the model consistently ignores the single most important constraint. Which combination of fixes is MOST likely to work?

    • A. Bury the constraint deeper and hope it is noticed.
    • B. Move the constraint to the top, label it clearly, and restate it as an explicit success criterion.
    • C. Remove the constraint to simplify.
    • D. Add ten unrelated constraints.
    Show answer

    Answer: B.

    Prominence, labelling and turning it into a checkable criterion all raise the chance the constraint is honoured. A reduces prominence. C abandons the requirement. D drowns it in noise.

  30. Q30D4 · Prompting and InstructionsSelect one

    Two prompts differ only in that one adds a named audience and a concrete success criterion. The second consistently produces better output. What does this BEST illustrate?

    • A. Longer prompts are always better.
    • B. Relevant, specific context and checkable criteria improve steerability more than length alone.
    • C. Audience never affects output.
    • D. Only reasoning models respond to audience.
    Show answer

    Answer: B.

    The improvement comes from relevant specifics and a checkable criterion, not raw length. A confuses the mechanism with length. C denies the observed effect. D wrongly restricts the principle to one model class.

  31. Q31D4 · Prompting and InstructionsSelect one

    A user wants a screening email that invites a call and names one specific skill from the job post. In prompt anatomy, where do those two requirements belong?

    • A. In the role component.
    • B. In the task and constraints, as explicit required actions and content.
    • C. In the audience only.
    • D. In the format only.
    Show answer

    Answer: B.

    Required actions and specific content are task and constraint elements. A sets a persona, not required content. C names who it is for, not what it must contain. D defines shape, not the required actions and skill mention.

  32. Q32D4 · Prompting and InstructionsSelect one

    A first draft is 90% right but uses the wrong currency and one wrong date. What is the MOST disciplined iteration move?

    • A. Regenerate from a fresh, unrelated prompt.
    • B. Give one targeted correction naming the currency and the correct date, keeping everything else.
    • C. Re-send the same prompt unchanged.
    • D. Switch models.
    Show answer

    Answer: B.

    A precise, minimal correction preserves the 90% that works and fixes the two errors. A discards good work. C repeats the same errors. D changes the tool when a small edit suffices.

  33. Q33D4 · Prompting and InstructionsSelect two

    Select the TWO changes that would MOST improve a prompt whose output was on-topic but flat and aimed at the wrong reader.

    • A. Add a tone-and-voice instruction with a short style example.
    • B. Name the intended audience and their level of expertise.
    • C. Add more prohibitions about what not to write.
    • D. Increase the reasoning effort to maximum.
    • E. Re-send the identical prompt several times.
    Show answer

    Answer: A and B.

    A tone instruction lifts flat copy and naming the audience fixes the wrong-reader issue. C adds negatives without a target. D adds compute unrelated to tone or audience. E repeats the same shortfall.

  34. Q34D4 · Prompting and InstructionsSelect one

    You want consistent two-sentence summaries across 200 documents with identical structure. Which approach gives the MOST consistent shape at scale?

    • A. Ask nicely and hope the structure holds.
    • B. Provide an explicit template and a worked example, and state the exact sentence count and content of each sentence.
    • C. Use a different phrasing for each document.
    • D. Raise reasoning effort per document.
    Show answer

    Answer: B.

    A template, example and explicit structure produce consistent shape across many items. A leaves consistency to chance. C guarantees variation. D adds cost without enforcing structure.

  35. Q35D4 · Prompting and InstructionsSelect one

    Which is the STRONGEST success criterion for a quarterly board update, judged by how checkable it is?

    • A. Make it comprehensive.
    • B. Two pages max, three named risks each with an owner and a mitigation, and figures reconciling to the finance pack.
    • C. Sound confident.
    • D. Cover everything important.
    Show answer

    Answer: B.

    B is specific and verifiable across length, content and reconciliation. A, C and D are vague and cannot be checked, so the model cannot reliably satisfy them or know when it has.

  36. Q36D4 · Prompting and InstructionsSelect one

    A colleague pads prompts with adjectives believing it raises quality, yet results are no better. What is the MOST accurate explanation to give?

    • A. Adjectives always improve quality; add more.
    • B. Decorative words add little; concrete role, context, constraints, format and a success criterion drive quality.
    • C. Quality depends only on model choice, never the prompt.
    • D. Prompts must avoid all adjectives.
    Show answer

    Answer: B.

    Substantive components, not decoration, drive quality. A is false. C ignores the well-established effect of prompt structure. D overcorrects into an absolute rule that is not the point.

  37. Q37D5 · Context, Files and MemorySelect two

    Using an information-placement view, which TWO statements about a reference document reused across many chats in a single workstream are correct?

    • A. It belongs in the workstream's Project so all its chats can draw on it.
    • B. Pasting it into every prompt is repetitive and error-prone.
    • C. It should be stored as a memory entry.
    • D. It should be placed in a scheduled task.
    • E. It cannot be reused and must be re-uploaded each chat individually.
    Show answer

    Answer: A and B.

    A reused reference document belongs in the Project, and pasting it repeatedly is inefficient and error-prone. C misuses memory for a document. D automates timing, not document access. E denies the reuse a Project explicitly provides.

  38. Q38D5 · Context, Files and MemorySelect one

    A user wants a small standing preference applied to all future chats without re-typing it, and asks how memory behaves. Which statement is MOST accurate?

    • A. Memory stores entire documents and large datasets by default.
    • B. Memory is suited to small, stable preferences, persists across chats, and can be viewed and deleted by the user.
    • C. Memory retrains the model on your data.
    • D. Memory only lasts for the current chat.
    Show answer

    Answer: B.

    Memory holds small stable preferences across chats and is user-manageable. A overstates its scope. C confuses memory with training. D contradicts its cross-chat persistence.

  39. Q39D5 · Context, Files and MemorySelect one

    A marathon chat kept open so nothing is lost is producing vaguer, slower answers. What is happening and the BEST fix?

    • A. The model is failing; nothing can be done.
    • B. The overloaded, aging context is diluting focus; extract the essentials into a summary or Project and start a fresh, focused chat.
    • C. The knowledge cutoff shifted.
    • D. Memory is full.
    Show answer

    Answer: B.

    A bloated, aging context degrades quality; summarising essentials and restarting restores focus. A gives up unnecessarily. C confuses training cutoff with context. D invents a memory-capacity cause.

  40. Q40D5 · Context, Files and MemorySelect one

    A one-off constraint that applies only to the current deliverable keeps getting saved to memory, so it wrongly affects later unrelated chats. What is the fix and the principle?

    • A. Keep it in memory; more persistence is always better.
    • B. Put one-off constraints in the immediate prompt, not memory; memory is for stable, recurring preferences.
    • C. Store it in a Project instead so it applies everywhere.
    • D. Delete all memory to be safe.
    Show answer

    Answer: B.

    One-off instructions belong in the prompt; only stable preferences belong in memory. A causes exactly the reported bleed-through. C would still over-apply it across the Project's chats. D is an overreaction that discards useful stable preferences.

  41. Q41D5 · Context, Files and MemorySelect one

    A consultant handles three clients and needs each client's preferences and files kept strictly separate with the LEAST ongoing effort. What is the BEST structure?

    • A. One shared chat with careful labelling.
    • B. A separate Project per client, each with its own instructions and files.
    • C. All three clients combined in one memory entry.
    • D. A new chat each time with everything pasted in.
    Show answer

    Answer: B.

    A Project per client keeps contexts cleanly separated with little ongoing effort. A risks cross-contamination. C mixes clients in memory. D is high-effort and error-prone every session.

  42. Q42D5 · Context, Files and MemorySelect two

    Select the TWO practices that BEST maintain clean context across a week of varied work.

    • A. Open a fresh chat per distinct task or topic.
    • B. Keep standing instructions in Projects and small stable preferences in memory rather than re-pasting them.
    • C. Attach every file you have ever used to each chat.
    • D. Keep one endless chat for all topics.
    • E. Store confidential client data in memory for speed.
    Show answer

    Answer: A and B.

    Fresh chats per topic and using Projects and memory appropriately keep context clean. C clutters every chat. D causes drift and dilution. E stores sensitive data where it should not live.

  43. Q43D6 · Verifying and Evaluating OutputSelect one

    A board pack is due in an hour; it reads fluently, but two subtotal rows do not match the headline total and one citation is unfindable. What should you do FIRST?

    • A. Submit it; it reads well and time is short.
    • B. Recompute the totals and resolve or remove the unfindable citation before the pack goes out.
    • C. Ask the model whether the pack is correct.
    • D. Change the layout to hide the mismatch.
    Show answer

    Answer: B.

    Internal-consistency and citation failures are blocking issues; fix the numbers and the source first. A ships known errors. C trusts the fallible source. D conceals the problem instead of resolving it.

  44. Q44D6 · Verifying and Evaluating OutputSelect one

    An analyst validates a summary's claim that a contract's notice period is 60 days. The contract was uploaded. What MOST appropriately validates it?

    • A. Trust the summary because the file was provided.
    • B. Open the contract and read the notice clause to confirm the 60-day figure.
    • C. Ask the model to re-read and confirm.
    • D. Compare it to a different contract.
    Show answer

    Answer: B.

    Grounding the claim means reading the actual clause in the source document. A over-trusts a summary that can misread. C re-uses the fallible generator. D compares to an irrelevant document rather than the source.

  45. Q45D6 · Verifying and Evaluating OutputSelect one

    A dashboard reports 92% accuracy on an AI-assisted tagging task and a manager wants to scale it. What matters MOST to check FIRST before expanding?

    • A. The overall percentage alone is enough to expand.
    • B. How errors are distributed and whether the costly or high-impact cases are the ones being missed.
    • C. Whether the interface colours are consistent.
    • D. Whether the model sounds confident.
    Show answer

    Answer: B.

    Aggregate accuracy can hide costly error patterns, so error distribution and impact matter most first. A ignores where the 8% falls. C is cosmetic. D relies on confidence, which is not an accuracy signal.

  46. Q46D6 · Verifying and Evaluating OutputSelect one

    For a number the model COMPUTED from data you provided, which validation action correctly matches its provenance?

    • A. Search the web to confirm the number.
    • B. Recompute it independently, for example in a spreadsheet or with a data-analysis tool.
    • C. Check a citation for it.
    • D. Ask a second chatbot.
    Show answer

    Answer: B.

    A computed figure is validated by independent recomputation, matching its provenance. A suits externally sourced facts, not computed ones. C applies to cited claims. D uses another fallible generator rather than recomputing.

  47. Q47D6 · Verifying and Evaluating OutputSelect one

    A researcher has a summary with eight citations and limited time. What is the MOST effective triage to catch fabrication FIRST?

    • A. Trust all eight because there are so many.
    • B. Spot-check the most load-bearing citations for existence and content, and treat any fabrication as a signal to check all of them.
    • C. Check the formatting of the citation list.
    • D. Ask the model to grade its own citations.
    Show answer

    Answer: B.

    Checking the claims that matter most, then escalating if any is fabricated, is efficient triage. A assumes quantity implies reliability. C checks style, not truth. D relies on the source that produced them.

  48. Q48D6 · Verifying and Evaluating OutputSelect two

    Select the TWO verification steps that are FREE and require NO external source.

    • A. Check that the answer actually addresses the question asked.
    • B. Check that the output's own numbers are internally consistent with its stated totals.
    • C. Confirm a statute against the official register.
    • D. Verify a market figure against an industry report.
    • E. Confirm a citation exists in a scholarly database.
    Show answer

    Answer: A and B.

    Relevance and internal-consistency checks need only the output itself. C, D and E each require an external authoritative source and so are neither free nor source-independent.

  49. Q49D6 · Verifying and Evaluating OutputSelect two

    Select the TWO practices that are weak checks and do NOT actually verify accuracy.

    • A. Asking the model whether it is confident in its answer.
    • B. Asking a second AI model and accepting the answer when both agree.
    • C. Reading the primary source to confirm the claim.
    • D. Recomputing a total in a spreadsheet.
    • E. Locating and opening the cited reference.
    Show answer

    Answer: A and B.

    Self-reported confidence and inter-model agreement are not grounding and can be confidently wrong. C, D and E are genuine verification against a source or by independent computation, so they are strong checks, not weak ones.

  50. Q50D6 · Verifying and Evaluating OutputSelect one

    A colleague argues that an internal-only forecast does not need verification because it will not be published. The forecast will set next year's budget. What is the BEST response?

    • A. Agree; internal outputs are exempt from checking.
    • B. Verify it, because the consequence of setting a budget, not whether it is published, determines the required rigour.
    • C. Verify only if someone senior asks.
    • D. Ask the model to re-run the forecast.
    Show answer

    Answer: B.

    Rigour scales with consequence, and a budget-setting forecast is high-stakes regardless of publication. A and C wrongly tie checking to audience or authority. D re-uses the same fallible generator instead of verifying.

  51. Q51D6 · Verifying and Evaluating OutputSelect one

    An output you will act on is long and reads smoothly. Applying a cost-ordered verification approach, which check comes FIRST?

    • A. A paid third-party audit.
    • B. A quick, free plausibility and internal-consistency read before any costlier external checks.
    • C. Commissioning a survey to confirm the claims.
    • D. Rewriting the whole document.
    Show answer

    Answer: B.

    Start with the cheapest, fastest check that catches obvious errors, then escalate only as needed. A and C are heavyweight later steps. D is not a verification step at all and wastes effort before checking correctness.

  52. Q52D7 · Responsible and Safe UseSelect one

    A company policy names an Enterprise workspace as the ONLY sanctioned place for confidential data, but a personal account is faster today. An employee is under deadline pressure. What should they do?

    • A. Use the personal account this once because it is faster.
    • B. Use only the sanctioned Enterprise workspace, because policy and data protection outrank convenience.
    • C. Split the data across both accounts.
    • D. Ask the model which account to use.
    Show answer

    Answer: B.

    Confidential data must stay in the sanctioned workspace regardless of deadline pressure. A violates policy and risks exposure. C still exposes data on an unsanctioned account. D delegates a governance decision to the model.

  53. Q53D7 · Responsible and Safe UseSelect one

    A vendor comparison the model produced lists many pros for Option A and mostly cons for Option B, matching how the user framed the request. What is the issue and the BEST fix?

    • A. No issue; the model is objective.
    • B. Framing bias amplified by sycophancy; re-request a balanced comparison with equal scrutiny of both options and the strongest case for B.
    • C. A hallucination; add citations.
    • D. A context-window overflow; start a new chat.
    Show answer

    Answer: B.

    The lopsided result reflects biased framing plus sycophancy; a neutral, balanced re-request fixes it. A ignores the bias. C addresses fabricated facts, not skew. D addresses long-chat drift, which is not the cause here.

  54. Q54D7 · Responsible and Safe UseSelect one

    When is disclosure of AI assistance MOST clearly expected?

    • A. When brainstorming privately for yourself.
    • B. When submitting work in a context whose rules or norms require disclosing AI use, such as certain academic or regulated deliverables.
    • C. Never; disclosure is always optional.
    • D. Only when the output is wrong.
    Show answer

    Answer: B.

    Disclosure is expected where the context's rules or norms require it, such as academic or regulated settings. A is private and low-stakes. C ignores real disclosure norms. D ties disclosure to correctness, which is not the criterion.

  55. Q55D7 · Responsible and Safe UseSelect one

    A model output defaults to gendered role assumptions, casting the engineer as male and the assistant as female. What is this and the correct FIRST response?

    • A. A harmless stylistic choice to leave as is.
    • B. A bias reflecting patterns in training data; correct it, use neutral phrasing, and review outputs for similar bias.
    • C. A hallucination requiring a citation.
    • D. A context-window problem.
    Show answer

    Answer: B.

    Default gendered roles are a form of bias to correct and watch for. A dismisses a real fairness issue. C misclassifies bias as fabrication. D is unrelated to how the bias arises.

  56. Q56D7 · Responsible and Safe UseSelect one

    An educator drafts a quiz with ChatGPT to use with students next week. What responsible-use step is MOST essential before using it?

    • A. Use it immediately; drafts are always accurate.
    • B. Review and verify the questions and answers for correctness, bias and appropriateness before using them with students.
    • C. Ask the model to confirm the quiz is correct.
    • D. Publish it and correct errors afterwards.
    Show answer

    Answer: B.

    AI-drafted educational content must be reviewed for accuracy, bias and fit before use with students. A over-trusts a draft. C relies on the same fallible source. D exposes students to unreviewed errors first.

  57. Q57D7 · Responsible and Safe UseSelect two

    Select the TWO outputs that MOST clearly require a mandatory human review gate before acting.

    • A. A published customer-facing legal disclaimer.
    • B. A recommendation that would reject job candidates.
    • C. A private list of brainstorming ideas for yourself.
    • D. A draft caption for a personal photo.
    • E. A set of synonyms for a word.
    Show answer

    Answer: A and B.

    A published legal disclaimer and a hiring rejection are high-stakes and consequential, needing human review. C, D and E are low-stakes and easily reversible, so they do not require a formal gate.

  58. Q58D2 · How Language Models BehaveSelect one

    A user insists the model must be right because it answered instantly and never said it was unsure. What is the MOST accurate response?

    • A. Agreed; speed and the absence of hedging prove correctness.
    • B. Neither speed nor an absence of hedging indicates accuracy; the model can be fast, confident and wrong, so the claim still needs verification.
    • C. The model always warns when it is unsure.
    • D. Instant answers come from a verified database.
    Show answer

    Answer: B.

    Confidence and speed are not accuracy signals; the model can be fluently wrong without hedging. A treats surface cues as proof. C is false; models often do not flag uncertainty. D invents a verified-lookup source that does not exist.

  59. Q59D3 · ChatGPT Surfaces and FeaturesSelect one

    A user needs a reusable team assistant that always applies a fixed review checklist AND can be shared, while a separate one-off question needs today's exchange rate. Which pairing is correct?

    • A. A custom GPT for the reusable assistant, and Search for the exchange rate.
    • B. Memory for the assistant, and image generation for the exchange rate.
    • C. A scheduled task for the assistant, and deep research for the exchange rate.
    • D. Canvas for both.
    Show answer

    Answer: A.

    A custom GPT packages a reusable, shareable assistant and Search fetches a fresh rate. B misuses memory and image generation. C uses a scheduler and heavyweight research for the wrong needs. D is an editing surface for neither task.

  60. Q60D5 · Context, Files and MemorySelect one

    A user needs a small, stable personal preference applied everywhere, a single 25-page report summarised once, and a client's standing brand rules shared with a team. Which mapping is correct?

    • A. Memory for the preference, a file upload for the report, and a shared Project for the brand rules.
    • B. A Project for the preference, memory for the report, and a scheduled task for the brand rules.
    • C. Memory for all three.
    • D. A single long chat for all three.
    Show answer

    Answer: A.

    Small stable preferences fit memory, a one-time document fits a file upload, and shared standing rules fit a shared Project. B misplaces each item. C overloads memory with a document and shared rules. D relies on a degrading single chat.

Last updated Sep 18, 2026