AI Cert Prep
Type to search documentation.

API Developer Path

API Developer Path – Mock Exam 2

A harder 60-item independent mock exam for the OpenAI API developer path, with multi-constraint and judgment items, full explanations and a readiness readout.

This is the second full-length independent mock exam for the API developer path, built from publicly available OpenAI learning objectives and product documentation. It is not an official OpenAI assessment: Academy badges and pathway certificates are not certifications. Mock 2 is deliberately harder than Mock Exam 1 — more multi-constraint stems, more FIRST / BEST / MOST cost-effective / TWO qualifiers, and more scenario framing. Treat it as your readiness gate: all 60 items are new and distinct from Mock 1 and the domain pages.

Instructions

  • Time: 90 minutes for 60 items — the pace our mock uses; OpenAI does not publish a public exam length.
  • Items: 60, single-answer and multiple-response. Each item states how many answers to select.
  • Selection: for a Select two item you must pick both correct options and no incorrect one to earn the mark. There is no guessing penalty, so answer every question.
  • Target: aim for at least 80% raw (about 48 of 60) here before you sit the real Academy assessment, which passes at the same 80% line. Clearing this harder mock at 80%+ is a strong sign of readiness.
  • Work each question before expanding the answer.

Domain distribution

#DomainItems here
D1Scoping AI Solutions7
D2The Responses API and Model Selection11
D3Evaluating AI Applications10
D4Designing and Building Agentic Systems11
D5Retrieval-Augmented Generation9
D6Performance, Latency and Cost7
D7Production Safety and Operations5

Total: 7 + 11 + 10 + 11 + 9 + 7 + 5 = 60 items.

Readiness interpretation

This is an independent readiness indicator, not an official score.

Raw score (of 60)BandInterpretation
54–60 (90%+)Strong readinessHandling the hardest multi-constraint items well
48–53 (80–89%)Assessment readyAt or above the 80% Academy line on the harder set
42–47 (70–79%)Building confidenceClose; drill the judgment and cost-arithmetic items
under 42 (less than 70%)Keep learningRevisit the domain pages and retake Mock 1 first

Take the mock exam

Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

60 questions · one at a time · 90-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · Scoping AI SolutionsSelect one

    A VP wants a fully autonomous incident-handling system live in one week, with no labelled incidents and a requirement that no customer-facing action fire without sign-off. What is the MOST appropriate FIRST step?

    • A. Scope a narrow assisted slice (draft responses with human approval) with a measurable metric, and stage autonomy only after evals justify it
    • B. Ship the full autonomous system and gate it later
    • C. Fine-tune on synthetic incidents to reach autonomy immediately
    • D. Choose gpt-6-astra and let capability compensate for the missing data
    Show answer

    Answer: A.

    With no labelled data and a hard human-sign-off requirement, the responsible first step is a narrow assisted slice with a metric, earning autonomy through evals. B contradicts the sign-off requirement and ships risk. C manufactures data that may not reflect reality and still skips evaluation. D leans on model size to paper over an absence of data and metrics, which capability cannot fix.

  2. Q2D1 · Scoping AI SolutionsSelect two

    You are scoping a feature and want to elicit only requirements that would CHANGE the architecture. Which TWO qualify?

    • A. Whether outputs feed an irreversible action that needs a human gate
    • B. Whether the workload is interactive or a bulk backlog
    • C. The company's preferred brand colours
    • D. Which font the report uses
    • E. The name of the project channel in chat
    Show answer

    Answer: A and B.

    Irreversibility drives approval gates and staging, and interactive-versus-bulk drives model, tier and streaming choices, so both change the architecture. C, D and E are cosmetic or organisational details that do not affect how the system is built.

  3. Q3D1 · Scoping AI SolutionsSelect one

    Two candidate metrics are proposed for a triage classifier: (i) 'users are happier' and (ii) 'macro-F1 ≥ 0.85 on a held-out labelled set with misroute rate below 3%'. Which is the BETTER success metric and why?

    • A. (ii), because it is specific, measurable against labelled data, and ties to a business-relevant error rate
    • B. (i), because user happiness is the ultimate goal
    • C. Neither; classifiers cannot be measured
    • D. Both are equivalent
    Show answer

    Answer: A.

    A good metric is specific and measurable against ground truth with a business-relevant threshold, which (ii) provides. (i) is unmeasurable as stated and cannot gate a release. C is false. D ignores that only (ii) can actually be computed and acted on.

  4. Q4D1 · Scoping AI SolutionsSelect one

    A workload does 90,000 classifications per day at 800 input and 120 output tokens. On gpt-5.6-luna ($0.20/$1.20 per M) it passes eval; a director prefers gpt-5.6-terra ($2/$12 per M). Roughly how much MORE per day does Terra cost for the same workload?

    • A. About $247 more per day (roughly $27 on Luna vs $274 on Terra)
    • B. About $25 more per day
    • C. Nothing; the prices are the same
    • D. About $2,470 more per day
    Show answer

    Answer: A.

    Daily input 90,000 x 800 = 72M tokens; output 90,000 x 120 = 10.8M tokens. Luna: 72 x $0.20 + 10.8 x $1.20 = $14.40 + $12.96 = ~$27. Terra: 72 x $2 + 10.8 x $12 = $144 + $129.60 = ~$274. The gap is about $247/day, so A. B is roughly the Luna total, not the gap; C is false since Terra is ten times pricier per token; D is ten times too high.

  5. Q5D1 · Scoping AI SolutionsSelect one

    During scoping you find the task's output auto-cancels shipments, an irreversible action. How should the solution plan BEST reflect this?

    • A. Record it as a hard constraint requiring an approval gate or reversible staging before any cancellation commits
    • B. Note it as a nice-to-have to revisit after launch
    • C. Rely on a higher-capability model to avoid mistakes
    • D. Log cancellations after the fact for auditing
    Show answer

    Answer: A.

    Irreversibility is a hard constraint that must shape the design with an approval gate or reversible staging before commit. B defers a safety-critical requirement. C cannot make an irreversible action safe by capability alone. D only records damage after it happens, which does not prevent it.

  6. Q6D1 · Scoping AI SolutionsSelect one

    A commodity capability with a mature hosted product exists, your team is small, and speed-to-value matters, but the data is highly sensitive and must never leave your network. What does this MOST likely change about the buy decision?

    • A. It shifts weight toward build or a deployment that keeps data in-network, because the data-residency constraint can outweigh the commodity argument
    • B. Nothing; always buy commodities
    • C. It means you must fine-tune your own model regardless
    • D. It means quality no longer matters
    Show answer

    Answer: A.

    A hard data-residency constraint can override the usual buy-the-commodity heuristic, favouring build or an in-network deployment. B ignores the constraint that changes the calculus. C leaps to fine-tuning, which residency does not require. D is a non sequitur.

  7. Q7D1 · Scoping AI SolutionsSelect one

    Which item belongs specifically in the 'Success metric' line of a solution plan, rather than the requirements or risks lines?

    • A. Macro-F1 at or above 0.85 on a held-out set, measured weekly
    • B. The data must stay in the EU
    • C. A prompt-injection attack could exfiltrate data
    • D. The feature must respond within two seconds
    Show answer

    Answer: A.

    A success metric states the measurable quality bar and how it is measured, which A does. B is a requirement/constraint, C is a risk, and D is a latency requirement, none of which is the quality success metric.

  8. Q8D2 · The Responses API and Model SelectionSelect one

    A pipeline needs strict schema output AND a live tool call AND to persist conversation state across turns, all on the primary interface. Which single API supports all three directly?

    • A. The Responses API, which supports structured outputs, function calling and server-side conversation state
    • B. The Chat Completions API, the recommended primary interface
    • C. The Assistants API, the current default
    • D. The Batch API
    Show answer

    Answer: A.

    The Responses API is the primary interface and supports structured outputs, function calling and conversation state together. B is legacy for new work. C is grouped under Legacy APIs. D is an offline processing tier, not an interactive interface with these features.

  9. Q9D2 · The Responses API and Model SelectionSelect two

    A scenario states: 'the model must return today's exchange rate and the customer's current balance'. Which TWO signals point toward tool calls rather than model recall?

    • A. The exchange rate changes constantly and is not in the model's training data
    • B. The customer balance is live private state the model cannot know
    • C. The request is phrased politely
    • D. The prompt is long
    • E. A larger model is being used
    Show answer

    Answer: A and B.

    Constantly changing external data and live private state are both outside model knowledge and must come from tools. C, D and E are surface features that say nothing about whether the needed information exists inside the model.

  10. Q10D2 · The Responses API and Model SelectionSelect one

    A team is on gpt-5.6-sol at high effort. Latency is unacceptable and their eval passes at medium. What is the MOST cost-effective correct change that preserves quality?

    • A. Drop to medium effort, since it passes the eval and reduces latency and cost
    • B. Upgrade to gpt-6-astra at xhigh
    • C. Keep high effort but add subagents
    • D. Raise temperature to speed generation
    Show answer

    Answer: A.

    The lowest effort that passes the eval is the target, and medium passes while cutting latency and cost. B increases both cost and latency. C adds parallel work unrelated to per-call effort latency. D does not reduce reasoning latency and harms determinism.

  11. Q11D2 · The Responses API and Model SelectionSelect one

    An application parses output_text and intermittently crashes when the model emits a tool call before any message. What is the ROOT fix, not a workaround?

    • A. Iterate the typed output items and handle tool-call items and message items explicitly rather than assuming a text message is always first
    • B. Wrap the parse in a try/except and ignore failures
    • C. Disable tools so no tool-call items appear
    • D. Retry the request until a text-first response occurs
    Show answer

    Answer: A.

    The root cause is treating the output as text-first when it is a sequence of typed items; handling each item type fixes it structurally. B hides errors and loses tool results. C removes a needed capability. D is a flaky workaround that wastes calls.

  12. Q12D2 · The Responses API and Model SelectionSelect one

    A conversation is approaching the 1.05M-token context window mid-task and must keep going. What is the appropriate action?

    • A. Apply compaction to summarise older context while retaining what the task still needs
    • B. Truncate the oldest messages blindly and hope nothing important is lost
    • C. Switch to a model with a smaller window
    • D. Raise max output tokens
    Show answer

    Answer: A.

    Compaction summarises stale context so the task continues within the window while preserving needed information. B risks discarding essential context. C makes the problem worse. D concerns output length, not the input context pressure.

  13. Q13D2 · The Responses API and Model SelectionSelect one

    A task is complex, open-ended and the single highest-value workflow in the product, requiring sustained multi-tool reasoning and judgment, at low volume. Which model is the BEST fit?

    • A. gpt-6-astra, built for the hardest end-to-end sustained-reasoning, multi-tool work
    • B. gpt-5.6-luna, to minimise cost
    • C. gpt-5.6-terra, the pragmatic all-rounder
    • D. gpt-5.5, the previous generation
    Show answer

    Answer: A.

    Astra is positioned for the hardest sustained-reasoning, judgment and multi-tool work, and low volume means its price is tolerable. B is for clear high-volume tasks, C is a mid all-rounder likely to fall short on the hardest work, and D steps back a generation.

  14. Q14D2 · The Responses API and Model SelectionSelect one

    A long-running research agent must survive well beyond one HTTP request AND deliver incremental progress to a UI. Which combination is correct?

    • A. Background mode for the detached run plus streaming (or webhooks) for progress
    • B. Prompt caching plus predicted outputs
    • C. A larger context window plus fine-tuning
    • D. Higher effort plus more output tokens
    Show answer

    Answer: A.

    Background mode detaches the run and streaming or webhooks carry incremental progress, matching both needs. B optimises cost and edit speed, not run lifetime. C and D address capacity and length, not detachment or progress reporting.

  15. Q15D2 · The Responses API and Model SelectionSelect one

    Which statement about the GPT-5.6 family is accurate?

    • A. Sol, Terra and Luna share a 1.05M-token context and 128K max output, and Astra is the priciest at $10/$50 per M
    • B. Luna has the largest context window of the family
    • C. Terra is more expensive than Sol
    • D. Astra is cheaper than Terra
    Show answer

    Answer: A.

    The lineup shows all four with 1.05M context and 128K max output, with Astra the most expensive at $10/$50. B is wrong: they share the same context. C is false: Terra ($2/$12) is cheaper than Sol ($4/$20). D is false: Astra is the most expensive, not cheaper than Terra.

  16. Q16D2 · The Responses API and Model SelectionSelect one

    A team keeps a two-turn conversation by re-sending the entire first turn's text on every follow-up 'to be safe', and their input costs are climbing. What is the BEST correction?

    • A. Use server-side conversation state so the follow-up references the prior response without resending its text
    • B. Switch to a cheaper model and keep resending everything
    • C. Lower reasoning effort to offset the resend cost
    • D. Accept the cost as unavoidable
    Show answer

    Answer: A.

    Conversation state removes the need to resend prior text, directly cutting the growing input cost. B treats a symptom while keeping the wasteful pattern. C does not address input tokens. D wrongly assumes the cost cannot be avoided.

  17. Q17D2 · The Responses API and Model SelectionSelect one

    You must choose the pragmatic all-rounder that is the natural replacement for existing GPT-5.5 workloads at moderate cost. Which model is it?

    • A. gpt-5.6-terra
    • B. gpt-6-astra
    • C. gpt-5.6-luna
    • D. gpt-5.5-pro
    Show answer

    Answer: A.

    Terra is described as the pragmatic all-rounder and the natural replacement for GPT-5.5 workloads. B is the top-tier flagship for the hardest work. C is for clear high-volume tasks. D is a costlier previous-generation model, not the replacement.

  18. Q18D2 · The Responses API and Model SelectionSelect one

    A guarantee is needed that a response is machine-parseable AND that a specific enum field only ever contains one of three allowed values. What MOST reliably enforces both?

    • A. A structured-output schema that constrains the field to the enum
    • B. A very detailed system prompt describing the enum
    • C. Post-hoc validation that rejects bad values and retries
    • D. Lowering the temperature only
    Show answer

    Answer: A.

    A schema with an enum constraint guarantees both parseability and the restricted value set. B is guidance the model can violate. C catches errors after the fact and wastes calls. D reduces variance but does not enforce an enum.

  19. Q19D3 · Evaluating AI ApplicationsSelect one

    After a model upgrade, the eval score drops on a cluster of long-input questions but is flat elsewhere. What is the MOST appropriate response?

    • A. Investigate the long-input cluster specifically, since the regression is localised, before deciding whether to adjust or roll back
    • B. Roll back immediately without diagnosis
    • C. Ignore it because overall score is fine
    • D. Delete the long-input questions from the dataset
    Show answer

    Answer: A.

    A localised regression should be diagnosed on the affected cluster to understand cause before acting. B discards a potential improvement without evidence. C ignores a real regression. D hides the problem by removing the very cases that reveal it.

  20. Q20D3 · Evaluating AI ApplicationsSelect two

    A prompt change fixes the single complaint you received, but you have no dataset. Before shipping, which TWO steps are the SAFEST?

    • A. Build a small representative eval set including the failing case and similar variations
    • B. Run the new and old prompts against that set and compare
    • C. Ship immediately since the one complaint is resolved
    • D. Increase the model tier to be safe
    • E. Ask the complainant to re-test in production
    Show answer

    Answer: A and B.

    Creating a small representative set with the failing case and comparing both prompts gives evidence the fix generalises without regressions. C ships on a single anecdote. D changes cost without evidence. E tests in production, exactly what an eval avoids.

  21. Q21D3 · Evaluating AI ApplicationsSelect one

    A grader for a support-triage classifier is being chosen. The task has known correct categories. Which grading approach is MOST appropriate and cheapest?

    • A. Deterministic label comparison against ground-truth categories
    • B. LLM-as-judge rating overall helpfulness
    • C. Human review of every prediction
    • D. Semantic similarity to a reference paragraph
    Show answer

    Answer: A.

    With known categories, comparing predicted to true labels is deterministic, precise and cheap. B introduces judge noise for a task with definite answers. C does not scale. D is meant for open text, not fixed categories.

  22. Q22D3 · Evaluating AI ApplicationsSelect one

    A team wants to trust an LLM-as-judge for faithfulness at scale. What is the correct ORDER of operations?

    • A. Label a sample by hand, measure judge-human agreement, then use the judge at scale only if agreement is adequate
    • B. Use the judge at scale, then check a few disagreements
    • C. Skip human labels and trust the largest model
    • D. Use the generator as its own judge
    Show answer

    Answer: A.

    Calibration precedes reliance: label a sample, verify agreement, then scale only if it is adequate. B scales an unvalidated judge first. C skips the calibration entirely. D risks self-preference bias with no independent check.

  23. Q23D3 · Evaluating AI ApplicationsSelect one

    Which is the STRONGEST reason online evaluation is needed even when offline evals pass?

    • A. Production traffic contains inputs and phrasings the fixed dataset never anticipated
    • B. Online evals are cheaper than offline
    • C. Offline evals cannot use graders
    • D. Online evals remove the need for a baseline
    Show answer

    Answer: A.

    Real traffic surfaces distributions the offline set could not foresee, which is precisely what online evaluation catches. B is not the reason and is not generally true. C is false. D is wrong: baselines still matter online.

  24. Q24D3 · Evaluating AI ApplicationsSelect one

    The prompt optimizer is proposed to improve a prompt. What does it REQUIRE to be genuinely useful?

    • A. An eval with examples and graders so improvements can be measured, not guessed
    • B. Only a single example prompt
    • C. A larger model
    • D. A production incident to trigger it
    Show answer

    Answer: A.

    Optimising a prompt meaningfully needs an eval with examples and graders to measure whether changes help. B is too little signal to optimise against. C is unrelated to optimisation feedback. D is not a prerequisite for optimisation.

  25. Q25D3 · Evaluating AI ApplicationsSelect one

    An eval dataset of 12 items gives scores that jump between 58% and 83% across identical runs. What is the MOST likely problem and fix?

    • A. The dataset is too small for a stable estimate; enlarge it to reduce variance
    • B. The model is nondeterministic and must be replaced
    • C. The graders should be removed
    • D. Reasoning effort must be maxed out
    Show answer

    Answer: A.

    On 12 items each flip moves the score by roughly eight points, so the swings are a small-sample artefact fixed by enlarging the set. B misattributes normal variance to a broken model. C removes the measurement. D does not address sample size.

  26. Q26D3 · Evaluating AI ApplicationsSelect one

    Given the current OpenAI docs, which statement about evals and fine-tuning is accurate?

    • A. The Evals API and fine-tuning are grouped under Legacy APIs, yet evals remain conceptually essential to shipping safely
    • B. Evals have been removed from the platform
    • C. Fine-tuning is the newest recommended surface
    • D. Evals are part of the Agents API core
    Show answer

    Answer: A.

    The docs group Evals and fine-tuning under Legacy APIs while evals still matter conceptually. B confuses legacy grouping with removal. C misstates fine-tuning's status. D misplaces evals inside the Agents API core.

  27. Q27D3 · Evaluating AI ApplicationsSelect one

    A grounded-answer eval passes at launch. Which change should trigger re-running it BEFORE shipping the change?

    • A. Any change to chunking, retrieval, prompt or model
    • B. A change to the UI colour theme
    • C. Renaming the project
    • D. Adding a new team member
    Show answer

    Answer: A.

    Grounding depends on chunking, retrieval, prompt and model, so any of those changes warrants re-running the eval. B, C and D do not touch the pipeline that determines whether answers stay supported by context.

  28. Q28D3 · Evaluating AI ApplicationsSelect one

    A team optimises cost by switching from Sol to Luna but does not re-run the eval. What is the BEST description of the risk?

    • A. Quality may have silently regressed below the bar because the change was never measured
    • B. Latency will necessarily increase
    • C. The context window shrinks below usable size
    • D. There is no risk; cheaper is always safe
    Show answer

    Answer: A.

    Changing the model without re-evaluating means any quality regression goes undetected until users feel it. B is not necessarily true and misses the point. C is false; the family shares 1.05M context. D ignores that cost cuts can trade away quality.

  29. Q29D4 · Designing and Building Agentic SystemsSelect one

    A startup needs the FASTEST path to a durable cloud agent with hosted sandboxed execution and long-running artifacts, and does not need EU residency or ZDR. Which runtime is the BEST fit?

    • A. The Agents API managed harness
    • B. The Agents SDK self-hosted
    • C. The Responses API with a hand-written loop
    • D. Chat Completions with cron
    Show answer

    Answer: A.

    The Agents API gives OpenAI-managed durable sessions, hosted sandboxes and artifacts with no self-hosting, and the US-only/no-ZDR constraint is acceptable here. B requires building and hosting the runtime. C and D lack managed durability, sandboxes and recovery.

  30. Q30D4 · Designing and Building Agentic SystemsSelect two

    Which TWO capabilities make the Agents SDK the right choice for a firm that must keep ALL processing under its own governance and infrastructure?

    • A. Self-hosted execution where the firm runs the agent loop
    • B. Code-first orchestration and guardrails in the firm's own environment
    • C. OpenAI-managed session recovery
    • D. Automatic hosted sandboxes billed as containers
    • E. A managed harness that runs the session for you
    Show answer

    Answer: A and B.

    Self-hosted execution and code-first orchestration/guardrails keep processing under the firm's control. C, D and E all describe the OpenAI-managed Agents API, which is the opposite of keeping everything under the firm's own governance.

  31. Q31D4 · Designing and Building Agentic SystemsSelect one

    An EU healthcare provider wants durable cloud agents but has a firm EU-residency and ZDR mandate. What is the correct conclusion about the Agents API, even with a self-hosted sandbox?

    • A. It is not eligible, because the Agents API is US-residency only with no ZDR and the sandbox choice does not change that
    • B. It is eligible if they self-host the sandbox in the EU
    • C. It is eligible with Private Link
    • D. It is eligible once mTLS is enabled
    Show answer

    Answer: A.

    The Agents API is US-residency only with no ZDR, and a self-hosted sandbox does not lift that; the mandate rules it out. B misreads the sandbox choice as changing residency. C and D address network security, not residency or retention.

  32. Q32D4 · Designing and Building Agentic SystemsSelect one

    In a managed agent session, why should you bound max_concurrent_subagents rather than leave it unbounded?

    • A. To cap parallel resource use, cost and blast radius of the delegated work
    • B. Because subagents cannot run in parallel at all
    • C. To increase the context window
    • D. To disable observability
    Show answer

    Answer: A.

    Bounding concurrency caps parallel resource use, cost and the blast radius of delegated work. B is false since subagents can run concurrently. C confuses concurrency with context size. D is unrelated and undesirable.

  33. Q33D4 · Designing and Building Agentic SystemsSelect one

    A prompt-injection payload hidden in a retrieved document tries to make an agent call a tool that exfiltrates data. Which control MOST directly limits the damage if the injection succeeds?

    • A. Least-privilege tool scopes so the agent cannot reach the exfiltration path
    • B. A larger model
    • C. Higher reasoning effort
    • D. A longer system prompt asking it to be careful
    Show answer

    Answer: A.

    If tools are scoped to least privilege, a hijacked agent simply cannot reach the exfiltration capability, containing the damage. B and C may reduce but cannot guarantee resistance to injection. D is guidance an injection can override.

  34. Q34D4 · Designing and Building Agentic SystemsSelect two

    You are building a triage-and-specialist agent system. Which TWO design elements correctly match their purpose?

    • A. A handoff transfers control from the triage agent to a billing agent for billing questions
    • B. A subagent is delegated a bounded research sub-task and returns its result to the parent
    • C. A handoff permanently deletes the triage agent
    • D. A subagent replaces the top-level orchestrator
    • E. Handoffs only work under the Batch API
    Show answer

    Answer: A and B.

    A handoff transfers control to a specialist, and a subagent handles a bounded sub-task and returns to the parent. C and D misdescribe the mechanics, and E places handoffs in an unrelated offline API.

  35. Q35D4 · Designing and Building Agentic SystemsSelect one

    A developer chooses the Agents API for a simple assistant that makes one internal function call, needs full per-step control, and has no session or sandbox requirement. Why is this the WRONG runtime?

    • A. It adds managed sessions and harness machinery the task does not need; the Responses API with function calling is simpler and sufficient
    • B. The Agents API cannot call functions
    • C. The Responses API cannot call tools
    • D. Function calling requires fine-tuning
    Show answer

    Answer: A.

    For a single tool call with full control and no session/sandbox needs, the managed harness is overkill; the Responses API with function calling is the right, simpler tool. B and C are false. D wrongly ties function calling to training.

  36. Q36D4 · Designing and Building Agentic SystemsSelect one

    An agent takes an occasional wrong action and the team cannot reconstruct the decision path afterward. Which capability BEST closes this gap?

    • A. Tracing and observability recording each step, tool call and its inputs and outputs
    • B. A bigger context window
    • C. More subagents for redundancy
    • D. A higher spend limit
    Show answer

    Answer: A.

    Reconstructing a past decision requires traces of steps, tool calls and their inputs/outputs. B, C and D add capacity, parallelism or budget but none of them produces the audit trail needed to explain a past action.

  37. Q37D4 · Designing and Building Agentic SystemsSelect one

    A team wants OpenAI to summarise context and resume sessions after interruptions for a multi-hour agent. Which runtime provides these as managed features?

    • A. The Agents API, which provides managed context summarisation and session resumption
    • B. The Agents SDK, which requires you to build them
    • C. The Responses API with a manual loop
    • D. The Batch API
    Show answer

    Answer: A.

    The Agents API's managed harness provides context summarisation and session resumption out of the box. B leaves these for you to implement. C requires you to hand-build durability. D is an offline bulk tier with none of these.

  38. Q38D4 · Designing and Building Agentic SystemsSelect one

    For an agent that can take irreversible actions, which combination is REQUIRED before it acts autonomously in production?

    • A. Least-privilege tool scopes plus a human approval gate on the irreversible step
    • B. A larger model plus higher effort
    • C. More subagents plus a bigger context window
    • D. A friendlier system prompt plus more output tokens
    Show answer

    Answer: A.

    Irreversible actions need least-privilege scoping to limit reach and a human gate to prevent unwanted commits. B, C and D adjust capability, parallelism or verbosity but none of them prevents an unwanted irreversible action.

  39. Q39D4 · Designing and Building Agentic SystemsSelect one

    How is a hosted-sandbox Agents API session billed compared with using a self-hosted sandbox, holding model and tools constant?

    • A. It adds container rates for the hosted sandbox on top of model and tool rates
    • B. It costs the same because the sandbox is free
    • C. It replaces model rates with a flat fee
    • D. It removes tool rates entirely
    Show answer

    Answer: A.

    Hosted sandboxes add container rates on top of the model and tool rates, which self-hosting avoids. B ignores the container charge. C and D misstate the billing model, which keeps model and tool rates and only adds the container cost.

  40. Q40D5 · Retrieval-Augmented GenerationSelect one

    Protocols are revised monthly and the previous version must NEVER be quoted. Re-indexing the new version is not enough. What else is REQUIRED?

    • A. Remove or supersede the old version in the index so retrieval can no longer return it
    • B. Ask the model not to quote old versions
    • C. Increase chunk overlap
    • D. Use a larger generation model
    Show answer

    Answer: A.

    If the old version remains retrievable, it can still be quoted, so it must be removed or superseded in the index. B relies on the model obeying and is bypassable. C and D affect retrieval granularity and generation, not whether stale content is present.

  41. Q41D5 · Retrieval-Augmented GenerationSelect one

    Retrieval recall is high (the right chunks are fetched) but faithfulness is low (answers add unsupported claims). What does this MOST likely indicate?

    • A. The generation step is not constrained to the retrieved context; tighten grounding instructions and eval for faithfulness
    • B. The retriever is broken
    • C. Chunks are too small
    • D. The embedding model must be replaced
    Show answer

    Answer: A.

    High recall with low faithfulness isolates the problem to generation not adhering to context, so grounding must be tightened and evaluated. B contradicts the high recall. C and D target retrieval, which is already working.

  42. Q42D5 · Retrieval-Augmented GenerationSelect two

    Which TWO are the STRONGEST reasons to prefer RAG over fine-tuning for a frequently changing internal knowledge task that must cite sources?

    • A. Content updates by re-indexing, without retraining
    • B. Answers can cite the retrieved source passages
    • C. Fine-tuning is always more accurate
    • D. RAG removes the need for any evaluation
    • E. RAG guarantees zero hallucination
    Show answer

    Answer: A and B.

    RAG handles frequent change by re-indexing and can cite retrieved sources, both decisive here. C is false as a blanket claim. D is wrong: RAG still needs grounding evals. E overpromises; grounding reduces but cannot guarantee zero hallucination.

  43. Q43D5 · Retrieval-Augmented GenerationSelect one

    A grounded assistant cites the wrong section for a correct-sounding answer. What is the BEST diagnostic step?

    • A. Inspect which chunks were retrieved and whether the cited chunk actually supports the claim
    • B. Immediately switch to a bigger model
    • C. Increase max output tokens
    • D. Turn off citations
    Show answer

    Answer: A.

    Diagnosing a mis-citation means examining the retrieved chunks and checking whether the cited one supports the claim, which localises retrieval versus generation fault. B and C change behaviour without diagnosis. D hides the very signal that exposes the problem.

  44. Q44D5 · Retrieval-Augmented GenerationSelect one

    Semantic retrieval works for prose but fails on queries containing exact SKUs like SKU-88231. What is the MOST direct remedy?

    • A. Add lexical/keyword matching to form a hybrid retriever
    • B. Retrain the embedding model on SKUs
    • C. Increase the number of retrieved chunks only
    • D. Lower the temperature
    Show answer

    Answer: A.

    Exact identifiers are matched by lexical search, so a hybrid retriever fixes the SKU misses directly. B is heavy and still meaning-biased. C fetches more chunks but not the right exact match. D affects generation, not retrieval.

  45. Q45D5 · Retrieval-Augmented GenerationSelect one

    A team validates RAG quality only by reading a handful of answers that 'sound right'. What is the BEST description of what is missing?

    • A. A grounded-answer eval measuring retrieval quality and faithfulness on a representative set
    • B. A larger vector store
    • C. More expensive embeddings
    • D. A bigger context window
    Show answer

    Answer: A.

    Sounding right is not measured grounding; a grounded-answer eval on a representative set is what is missing. B, C and D change infrastructure but none of them measures whether answers are actually supported by retrieved context.

  46. Q46D5 · Retrieval-Augmented GenerationSelect one

    Answers are bloated with unrelated text after chunk size was raised to 5,000 tokens. What is the BEST first adjustment?

    • A. Reduce chunk size to more focused passages so each retrieved chunk is on-topic
    • B. Increase chunk size further
    • C. Switch to a bigger generation model
    • D. Disable overlap
    Show answer

    Answer: A.

    Oversized chunks carry unrelated content, so shrinking them to focused passages reduces the noise. B worsens it. C changes generation, not the noisy retrieval input. D risks splitting boundary-spanning answers without fixing bloat.

  47. Q47D5 · Retrieval-Augmented GenerationSelect one

    A clinician-facing RAG system must restrict each user to their own department's protocols. A developer proposes enforcing this in the system prompt. Why is that insufficient?

    • A. Prompt instructions are bypassable; authorisation must filter documents in the retrieval layer so out-of-scope content is never fetched
    • B. Prompts are too short to hold the rule
    • C. The model cannot read system prompts
    • D. It would slow retrieval down
    Show answer

    Answer: A.

    Access control is a security boundary that must live in retrieval; a prompt rule can be circumvented and out-of-scope documents would still be retrievable. B and C are false. D is not the reason; correctness and security are.

  48. Q48D5 · Retrieval-Augmented GenerationSelect one

    A developer wants to stuff a whole 900-page manual into every prompt because it fits in 1.05M tokens. Which is the STRONGEST argument for RAG instead?

    • A. Sending only the relevant passages cuts per-request cost and latency and reduces distraction from irrelevant text
    • B. The manual does not fit in the window
    • C. Prompts cannot contain long documents
    • D. RAG guarantees perfect answers
    Show answer

    Answer: A.

    Even when it fits, paying to process the entire manual on every call is wasteful and can distract the model; RAG sends only what the query needs. B is false since it fits. C is untrue. D overpromises perfection.

  49. Q49D6 · Performance, Latency and CostSelect one

    A feature is both too slow and too expensive on gpt-5.6-sol at high effort. What is the correct ORDER of operations to fix it?

    • A. Measure where time and cost go, then lower effort/model to the least that passes the eval, then apply caching or batching where they fit
    • B. Immediately move everything to the Batch API
    • C. Switch to Luna and skip evaluation
    • D. Raise effort to reduce retries
    Show answer

    Answer: A.

    Measure first, then reduce effort/model to the cheapest that still passes the eval, then apply caching or batching where appropriate. B ignores interactivity needs and skips measurement. C changes model with no quality check. D adds latency and cost.

  50. Q50D6 · Performance, Latency and CostSelect two

    An eval shows gpt-5.6-luna meets the same target gpt-5.6-sol currently hits on a high-volume task. Which TWO actions are the MOST cost-effective and safe?

    • A. Switch the workload to Luna since it passes the eval
    • B. Keep monitoring the eval after the switch to catch any regression
    • C. Stay on Sol to be safe
    • D. Move to Astra for extra headroom
    • E. Turn evals off after switching
    Show answer

    Answer: A and B.

    Luna passes, so switching cuts cost, and continued monitoring guards against regression, which is both cost-effective and safe. C keeps paying more for no quality gain. D increases cost further. E removes the very check that makes the switch safe.

  51. Q51D6 · Performance, Latency and CostSelect one

    Which workload is the WRONG fit for the Batch API?

    • A. A synchronous chat feature where users wait for each reply
    • B. An overnight document-classification run
    • C. A large offline embedding job
    • D. A weekly bulk summarisation of archived tickets
    Show answer

    Answer: A.

    Batch trades latency for cost and is unsuitable for a synchronous feature where users wait in real time. B, C and D are latency-tolerant offline jobs that are exactly what Batch is for.

  52. Q52D6 · Performance, Latency and CostSelect one

    A workload runs 500,000 requests per day at 1,500 input and 500 output tokens on gpt-5.6-terra ($2/M input, $12/M output). Roughly what is the daily cost?

    • A. About $4,500 per day
    • B. About $450 per day
    • C. About $1,500 per day
    • D. About $45,000 per day
    Show answer

    Answer: A.

    Input: 500,000 x 1,500 = 750M x $2/M = $1,500. Output: 500,000 x 500 = 250M x $12/M = $3,000. Total ~$4,500/day, so A. B is a tenth too small, C omits the output cost, and D is ten times too high.

  53. Q53D6 · Performance, Latency and CostSelect one

    A latency-tolerant service wants cheaper-than-standard processing but cannot accept Batch's hours-long turnaround. Which tier is the BEST fit and why?

    • A. Flex processing, which lowers cost for latency-tolerant work without Batch's long delay
    • B. Fast mode, which prioritises speed at higher cost
    • C. The Realtime API, for interactive audio
    • D. Standard synchronous processing, the baseline cost
    Show answer

    Answer: A.

    Flex processing is the middle tier that reduces cost for latency-tolerant work without the multi-hour Batch wait. B optimises speed and costs more. C is for real-time interaction. D is the baseline the service is trying to beat on price.

  54. Q54D6 · Performance, Latency and CostSelect two

    Time-to-first-token is the dominant latency term for an interactive feature on a reasoning model at high effort. Which TWO levers help MOST directly?

    • A. Lower the reasoning effort to the least that still passes the eval
    • B. Cache the large shared prompt prefix so it is not reprocessed each call
    • C. Increase max output tokens
    • D. Add more retrieved context to every request
    • E. Move to the Batch API
    Show answer

    Answer: A and B.

    Lower effort shortens the thinking before the first token, and caching the shared prefix removes reprocessing time, both cutting time-to-first-token. C lengthens output. D adds input to process. E is offline and irrelevant to interactive latency.

  55. Q55D6 · Performance, Latency and CostSelect one

    A document-editing feature regenerates a 20,000-token document where roughly 97% is unchanged. Which feature cuts latency MOST, and why?

    • A. Predicted outputs, because most of the output is already known and can be accelerated
    • B. Prompt caching, because the input prefix repeats
    • C. The Batch API, because the job is large
    • D. A larger context window, because the document is long
    Show answer

    Answer: A.

    Predicted outputs speed generation when most of the output equals a known reference, exactly the near-identical edit case. B helps repeated inputs, not near-identical outputs. C is offline. D adds capacity but not generation speed for known output.

  56. Q56D7 · Production Safety and OperationsSelect one

    Under sustained load your service receives frequent 429 responses. Beyond exponential backoff, which control BEST prevents overrunning your allocation?

    • A. Client-side rate limiting aligned to your account rate limits
    • B. Retrying instantly without limit
    • C. Raising reasoning effort
    • D. Increasing max output tokens
    Show answer

    Answer: A.

    Shaping outbound traffic to your rate limits prevents repeatedly hitting 429s in the first place, complementing backoff. B intensifies the overload. C and D affect quality and length, not request rate.

  57. Q57D7 · Production Safety and OperationsSelect one

    A sensitive workload must reach the API WITHOUT traversing the public internet. Which control fits BEST?

    • A. Private Link for a private network path to the API
    • B. An IP allowlist
    • C. A spend limit
    • D. A higher rate limit
    Show answer

    Answer: A.

    Private Link provides a private network path that avoids the public internet. B restricts which public IPs may connect but still uses the public internet. C caps cost. D changes throughput, neither of which provides a private path.

  58. Q58D7 · Production Safety and OperationsSelect two

    You are writing a pre-launch checklist for an agent that can take irreversible actions on customer accounts. Which TWO items MUST be on it?

    • A. A human approval gate on every irreversible action
    • B. Least-privilege tool scopes and a spend limit
    • C. The maximum available context window
    • D. The most expensive model on every call
    • E. Disabling logging to reduce noise
    Show answer

    Answer: A and B.

    An approval gate stops unwanted irreversible actions and least-privilege scopes plus spend limits contain reach and cost. C and D raise capacity or cost without addressing safety. E removes the observability needed to investigate incidents.

  59. Q59D7 · Production Safety and OperationsSelect one

    What is the key difference between red teaming and misalignment monitoring?

    • A. Red teaming proactively probes for weaknesses before launch; misalignment monitoring watches a deployed system for behaviour drifting from intent
    • B. They are the same activity
    • C. Red teaming only runs in production; monitoring only runs pre-launch
    • D. Both are billing controls
    Show answer

    Answer: A.

    Red teaming is a proactive pre-launch probe for weaknesses, while misalignment monitoring is an ongoing watch on deployed behaviour. B denies a real distinction. C reverses their timing. D miscategorises safety practices as billing.

  60. Q60D7 · Production Safety and OperationsSelect one

    A read-only reporting job uses an API key that also carries permission to delete resources. Which principle is violated and what is the correct fix?

    • A. Least privilege; issue a scoped key with read-only permissions for the job
    • B. Idempotency; make the job retry-safe
    • C. Backpressure; throttle the job
    • D. Statelessness; remove server state
    Show answer

    Answer: A.

    A reporting job holding delete rights violates least privilege; the fix is a read-only scoped key. B, C and D are legitimate concepts but none describes granting more permissions than the task needs.

Last updated Sep 18, 2026