AI Cert Prep
Type to search documentation.

API Developer Path

API Developer Path – Mock Exam 1

A 60-item independent mock exam for the OpenAI API developer path, blueprint-weighted across seven domains, with full explanations and a readiness readout.

This is a full-length, blueprint-weighted independent mock exam for the API developer path built from publicly available OpenAI learning objectives and product documentation. It is not an official OpenAI assessment and not the Academy assessment: Academy badges and pathway certificates are not certifications. All 60 items are new and do not repeat the domain-page questions. Use this as your diagnostic sitting before the harder Mock Exam 2.

Instructions

  • Time: 90 minutes for 60 items — the pace our mock uses; OpenAI does not publish a public exam length.
  • Items: 60, single-answer and multiple-response. Each item states how many answers to select.
  • Selection: for a Select two item you must pick both correct options and no incorrect one to earn the mark. There is no guessing penalty, so answer every question.
  • Target: aim for at least 80% raw (about 48 of 60) before you sit the real Academy assessment, which passes at the same 80% line.
  • Work each question before expanding the answer.

Domain distribution

#DomainItems here
D1Scoping AI Solutions7
D2The Responses API and Model Selection11
D3Evaluating AI Applications10
D4Designing and Building Agentic Systems11
D5Retrieval-Augmented Generation9
D6Performance, Latency and Cost7
D7Production Safety and Operations5

Total: 7 + 11 + 10 + 11 + 9 + 7 + 5 = 60 items.

Readiness interpretation

This is an independent readiness indicator, not an official score.

Raw score (of 60)BandInterpretation
54–60 (90%+)Strong readinessConfident across all domains
48–53 (80–89%)Assessment readyAt or above the 80% Academy line; tidy up weak domains
42–47 (70–79%)Building confidenceClose; target the domains that cost you marks
under 42 (less than 70%)Keep learningRevisit the domain pages before re-attempting

Take the mock exam

Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

60 questions · one at a time · 90-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · Scoping AI SolutionsSelect one

    A stakeholder describes a goal as 'make the chatbot smarter'. Before writing any code, what turns this into a scopable problem?

    • A. Define the specific task, the user, the input, the expected output and a measurable success metric
    • B. Pick the most capable model so quality is never the bottleneck
    • C. Start collecting a fine-tuning dataset immediately
    • D. Build a prototype and iterate on stakeholder feedback later
    Show answer

    Answer: A.

    Scoping converts a vague wish into a task with defined input, output, user and a measurable success metric, which is what every later decision depends on. B jumps to a model before the problem is known and wastes budget. C presumes fine-tuning is needed before any task definition exists. D skips the definition that a prototype would be built to test, so feedback would be unfocused.

  2. Q2D1 · Scoping AI SolutionsSelect one

    Which of these is a deterministic requirement that should NOT rely on a language model as the source of truth?

    • A. Computing an account's exact outstanding balance from ledger rows
    • B. Summarising the tone of a customer email
    • C. Drafting a suggested reply for an agent to edit
    • D. Classifying a ticket into one of five categories
    Show answer

    Answer: A.

    An exact balance must come from arithmetic over authoritative data, not a model that can approximate; the model can present or explain it but not compute it as truth. B, C and D are judgment or language tasks where an approximate, reviewable output is acceptable, which is exactly where models fit.

  3. Q3D1 · Scoping AI SolutionsSelect one

    A team must decide between building a custom transcription pipeline and buying a hosted product for support-call transcription. Which factor MOST favours buying?

    • A. Transcription is a commodity capability with mature hosted options and no competitive differentiation
    • B. The team enjoys building infrastructure
    • C. They want to avoid all third-party APIs on principle
    • D. They have a large ML research group with spare time
    Show answer

    Answer: A.

    Buy when the capability is a commodity that provides no differentiation and mature products exist, so effort goes to the product instead. B is a preference, not a business reason. C is a constraint that would argue against buying, not for it. D suggests build capacity but capacity is not a reason to build something with no differentiation value.

  4. Q4D1 · Scoping AI SolutionsSelect one

    You are scoping a summariser for 40,000 documents per day at ~1,200 input and ~200 output tokens each; gpt-5.6-luna meets the quality bar. Luna costs $0.20 per million input tokens and $1.20 per million output tokens. Roughly what is the daily cost?

    • A. About $19 per day
    • B. About $2 per day
    • C. About $58 per day
    • D. About $190 per day
    Show answer

    Answer: A.

    Input: 40,000 x 1,200 = 48M tokens x $0.20/M = $9.60. Output: 40,000 x 200 = 8M tokens x $1.20/M = $9.60. Total ~$19.20/day, so A. B undercounts by an order of magnitude, C uses Terra-like rates, and D is ten times too high.

  5. Q5D1 · Scoping AI SolutionsSelect two

    A stakeholder wants an AI feature that auto-approves expense reports and posts them to the ledger with no human review. Which TWO points should shape your scope?

    • A. The posting step is irreversible, so it needs a human approval gate or reversible staging before commit
    • B. Ledger totals must be computed deterministically, not generated by the model
    • C. The most expensive model guarantees the numbers will be correct
    • D. No success metric is needed because the workflow is fully automated
    • E. Evals are unnecessary once the prompt looks good on a few examples
    Show answer

    Answer: A and B.

    A recognises an irreversible action needs a gate or reversible staging, and B keeps arithmetic deterministic - both reshape the design. C is false: a bigger model does not make generated arithmetic authoritative. D is wrong because automation raises, not removes, the need for a metric. E dismisses evaluation, which is exactly what a money-moving feature requires.

  6. Q6D1 · Scoping AI SolutionsSelect one

    A director wants gpt-6-astra mandated for every feature 'to be safe'. A high-volume extraction task passes your eval at 96% on gpt-5.6-luna. What is the BEST response?

    • A. Show the eval result and cost difference and recommend Luna for this task, reserving Astra for tasks that need it
    • B. Comply and run everything on Astra to avoid the conversation
    • C. Switch to a competitor to escape the mandate
    • D. Run both models on every request and compare
    Show answer

    Answer: A.

    Model selection is per task and driven by eval results against the quality bar and cost; presenting evidence lets the team choose Luna where it suffices. B wastes large multiples of cost with no quality benefit. C is disproportionate. D doubles cost and latency on every request for no production value.

  7. Q7D1 · Scoping AI SolutionsSelect one

    During scoping you learn a feature must respond while a user waits AND another must clear a nightly backlog of 3M items. What does correct scoping conclude?

    • A. They are two workloads with different latency needs and likely different processing tiers, and should be scoped separately
    • B. One model and one processing tier should serve both to keep it simple
    • C. Both should use the Batch API for cost
    • D. Both should use synchronous streaming to keep users happy
    Show answer

    Answer: A.

    Interactive and bulk workloads have opposite latency profiles, so scoping treats them separately and can pick Batch for the backlog and low-latency processing for the interactive path. B forces a single tier onto incompatible needs. C would make the interactive feature unusably slow. D would waste money and offer no benefit on the backlog.

  8. Q8D2 · The Responses API and Model SelectionSelect one

    For a new API integration in September 2026, which interface should you build on for the long term?

    • A. The Responses API, the primary interface for new work
    • B. The Chat Completions API, because it is newest
    • C. The Assistants API, now the recommended default
    • D. The Agent Builder, the primary interface
    Show answer

    Answer: A.

    The Responses API is the primary interface for new work; Chat Completions is legacy for new builds. B misstates Chat Completions as new. C and D are wrong: Assistants and Agent Builder are grouped under Legacy APIs, not the recommended default.

  9. Q9D2 · The Responses API and Model SelectionSelect one

    You need gpt-5.6-terra to run with the least thinking that still passes your eval. Which principle applies?

    • A. Use the lowest reasoning effort that gets the result
    • B. Always use max effort for correctness
    • C. Copy the exact effort setting you used on GPT-5.5
    • D. Reasoning effort has no impact on cost or latency
    Show answer

    Answer: A.

    The documented guidance is to use the lowest reasoning effort that gets the result, because higher effort adds latency and cost. B over-spends by default. C is wrong: there is no exact mapping from GPT-5.5 efforts to GPT-5.6. D is false since more reasoning generally means more tokens, latency and cost.

  10. Q10D2 · The Responses API and Model SelectionSelect one

    A classification service must emit {"queue": "...", "priority": 1} that downstream code parses without defensive cleanup. What is the BEST mechanism?

    • A. Request structured outputs bound to a schema so the shape is guaranteed
    • B. Ask the model politely to 'return valid JSON only' in the prompt
    • C. Post-process free text with regular expressions
    • D. Raise the temperature so the model is more creative with formatting
    Show answer

    Answer: A.

    Structured outputs bound to a schema guarantee the parseable shape. B is a request the model can still violate. C is brittle and fails on edge cases. D makes output less predictable, the opposite of what is needed.

  11. Q11D2 · The Responses API and Model SelectionSelect one

    An assistant must fetch a customer's live subscription tier from an internal billing service to answer. What is the correct mechanism?

    • A. Function calling: expose a tool the model can invoke to read the billing service
    • B. Rely on the model's training knowledge of the customer
    • C. Paste the entire billing database into the system prompt
    • D. Increase max output tokens so the model can recall more detail
    Show answer

    Answer: A.

    Live, per-customer data must be fetched at request time through a tool the model calls, i.e. function calling. B is impossible: the model has no knowledge of live account state. C is infeasible and a data-exposure risk. D confuses output length with data access.

  12. Q12D2 · The Responses API and Model SelectionSelect one

    A high-volume extraction task runs 3M times per day and gpt-5.6-luna passes your eval at 97%. Which model should you deploy?

    • A. gpt-5.6-luna, because it meets the bar at the lowest cost for clear high-volume work
    • B. gpt-6-astra, because more capability is always safer
    • C. gpt-5.6-sol, to leave headroom
    • D. gpt-5.5-pro, the previous flagship
    Show answer

    Answer: A.

    Luna is designed for clear, repeatable, high-volume tasks and it already meets the quality bar, so it is the cheapest correct choice. B, C and D all cost multiples more for a task Luna passes, with no quality justification, and D also picks a costlier previous-generation model.

  13. Q13D2 · The Responses API and Model SelectionSelect one

    You must maintain a two-turn conversation without resending the first turn's text on the second call. Which Responses API feature supports this?

    • A. Server-side conversation state that links the follow-up to the prior response
    • B. Setting a higher temperature on the second call
    • C. Manually concatenating all prior text every time is the only option
    • D. Fine-tuning the model on the first turn
    Show answer

    Answer: A.

    The Responses API supports server-side conversation state so a follow-up references the prior response without resending its text. B is unrelated to state. C is the thing conversation state exists to avoid. D is a training operation, not a way to carry a single conversation.

  14. Q14D2 · The Responses API and Model SelectionSelect one

    A run must continue after a single HTTP request ends and report progress via streaming or webhooks. Which Responses API feature fits?

    • A. Background mode, which lets a run outlive one request
    • B. Prompt caching
    • C. Predicted outputs
    • D. Higher max output tokens
    Show answer

    Answer: A.

    Background mode lets a long-running task continue beyond one HTTP request and be followed via streaming or webhooks. B reduces cost on repeated prefixes but does not detach the run. C speeds edits, unrelated to lifetime. D only affects response length.

  15. Q15D2 · The Responses API and Model SelectionSelect two

    Which TWO practices reduce input-token cost on a very long multi-turn conversation without dropping needed context?

    • A. Use compaction to summarise older turns while keeping recent detail
    • B. Rely on server-side conversation state instead of resending full history each call
    • C. Raise reasoning effort to max
    • D. Increase temperature
    • E. Switch every call to gpt-6-astra
    Show answer

    Answer: A and B.

    Compaction summarises stale history and conversation state avoids resending prior text, both cutting input tokens. C increases spend and latency. D changes randomness, not tokens. E raises the per-token price, worsening cost.

  16. Q16D2 · The Responses API and Model SelectionSelect one

    When you read a Responses API result from a reasoning model that also called a tool, why can indexing the first output item be unreliable for the answer text?

    • A. The output is a sequence of typed items that can include reasoning and tool-call items, so you should read the aggregated output text rather than assume item zero is the answer
    • B. The first item is always an error
    • C. Reasoning models never produce text output
    • D. Tool calls overwrite all text items
    Show answer

    Answer: A.

    A response is a list of typed items; reasoning and tool-call items can precede the message, so code should use the aggregated output text helper rather than a fixed index. B, C and D are false generalisations that would break in the common case where text and tool items coexist.

  17. Q17D2 · The Responses API and Model SelectionSelect one

    A complex, open-ended, high-value analysis fails on gpt-5.6-terra in your eval and volume is low. Which is the BEST next model to try?

    • A. gpt-5.6-sol, built for complex open-ended high-value work
    • B. gpt-5.6-luna, to save money
    • C. gpt-5.5, the previous generation
    • D. Stay on Terra and raise temperature
    Show answer

    Answer: A.

    Sol is positioned for complex, open-ended, high-value work and is the natural step up from Terra, especially at low volume where its price matters less. B moves down the capability ladder for a task Terra already failed. C steps back a generation. D changes randomness, not capability.

  18. Q18D2 · The Responses API and Model SelectionSelect one

    What is TRUE about reasoning effort levels on the GPT-5.6 family?

    • A. Sol, Terra and Luna support none through max; Astra adds xhigh
    • B. Only Astra supports reasoning effort
    • C. Effort settings map one-to-one from GPT-5.5
    • D. Effort cannot be set below high
    Show answer

    Answer: A.

    Per the model table, the GPT-5.6 models support none through max and Astra additionally offers xhigh. B is false since Sol, Terra and Luna also take effort. C is explicitly denied by the docs. D is invented; none and low are available.

  19. Q19D3 · Evaluating AI ApplicationsSelect one

    A developer edits a prompt, checks two examples that look better, and wants to ship. What should they do FIRST?

    • A. Run the change against a representative eval dataset and compare to the baseline
    • B. Deploy and watch for complaints
    • C. Ask a colleague if it looks fine
    • D. Raise the model tier to be safe
    Show answer

    Answer: A.

    Two examples cannot show whether a change helps overall; a representative eval compared to a baseline can. B ships blind. C is anecdotal. D changes cost without evidence the prompt change is even good.

  20. Q20D3 · Evaluating AI ApplicationsSelect one

    You must grade whether extracted invoice fields exactly match ground truth. Which grader is BEST?

    • A. An exact-match/string grader against the known correct values
    • B. An LLM-as-judge scoring general quality
    • C. Human review of every item forever
    • D. A semantic-similarity grader
    Show answer

    Answer: A.

    Exact field matching against ground truth is a deterministic check, so a string/exact-match grader is precise and cheap. B adds noise and cost for a task with a definite answer. C does not scale. D would pass near-misses that must be exact.

  21. Q21D3 · Evaluating AI ApplicationsSelect one

    What is the difference between offline and online evaluation?

    • A. Offline runs on a fixed dataset before deploy; online measures behaviour on real production traffic
    • B. Offline is for training only; online is for testing only
    • C. They are two names for the same thing
    • D. Online evaluation replaces the need for any dataset
    Show answer

    Answer: A.

    Offline evaluation uses a curated dataset before release; online evaluation observes live traffic to catch inputs the dataset never contained. B mislabels their purposes. C denies a real distinction. D is wrong: online complements, not replaces, offline datasets.

  22. Q22D3 · Evaluating AI ApplicationsSelect one

    Before trusting an LLM-as-judge grader that scores summaries for faithfulness, what must you do?

    • A. Validate the judge against human labels on a sample to confirm it agrees with people
    • B. Assume it is correct because it is a large model
    • C. Use the same model as both generator and judge without checking
    • D. Skip the dataset and grade in production only
    Show answer

    Answer: A.

    An automated judge is only trustworthy once it has been calibrated against human labels on a sample. B is unearned trust. C risks the judge favouring its own style with no verification. D removes the controlled comparison that calibration needs.

  23. Q23D3 · Evaluating AI ApplicationsSelect one

    A team runs only offline evals and keeps getting production failures on phrasings never in their dataset. What is missing?

    • A. Online evaluation and sampling of real traffic to expand the dataset
    • B. A more expensive model
    • C. A larger max output token setting
    • D. More reasoning effort
    Show answer

    Answer: A.

    Failures on unseen phrasings are exactly what online evaluation and traffic sampling catch, feeding new cases back into the dataset. B, C and D change model behaviour but do nothing to surface the inputs the offline set never covered.

  24. Q24D3 · Evaluating AI ApplicationsSelect two

    Which TWO are legitimate reasons to include human review in an evaluation process?

    • A. To calibrate an automated judge against human judgment on a sample
    • B. To adjudicate high-stakes or ambiguous cases the graders disagree on
    • C. To review every single production request forever
    • D. To replace the eval dataset entirely
    • E. To make the eval slower on purpose
    Show answer

    Answer: A and B.

    Human review calibrates judges and resolves high-stakes or ambiguous cases. C does not scale and is not the goal of review. D discards the reusable dataset humans help build. E is not a reason at all.

  25. Q25D3 · Evaluating AI ApplicationsSelect two

    Which TWO purposes does keeping a baseline eval score serve?

    • A. Detecting whether a prompt, model or pipeline change is an improvement
    • B. Detecting whether a change is a regression before it ships
    • C. Satisfying a required platform field
    • D. Making the dashboard look complete
    • E. Replacing the need for online evaluation
    Show answer

    Answer: A and B.

    A baseline is the reference that lets you see both improvements and regressions from a change. C invents a platform requirement. D is cosmetic. E is wrong: baselines and online evaluation address different things and one does not replace the other.

  26. Q26D3 · Evaluating AI ApplicationsSelect one

    How should you describe the Evals API given the current OpenAI docs structure?

    • A. The Evals API, now grouped with Legacy APIs, though evals still matter conceptually
    • B. The newest and recommended primary surface
    • C. A deprecated feature that no longer works
    • D. Part of the Responses API core
    Show answer

    Answer: A.

    The docs group Evals under Legacy APIs while evals remain conceptually important, so A is the accurate framing. B overstates its status. C is wrong: legacy is not the same as removed. D misplaces it in the Responses core.

  27. Q27D3 · Evaluating AI ApplicationsSelect one

    Your eval dataset has 10 items and scores swing widely between runs. What is the MOST likely problem?

    • A. The dataset is too small to give a stable estimate
    • B. The model is broken
    • C. The graders are all wrong
    • D. Reasoning effort is too low
    Show answer

    Answer: A.

    With only 10 items, a single flip moves the percentage sharply, so the instability is a sample-size problem; enlarge the set. B, C and D are possible in general but the described symptom - large swings on a tiny set - points squarely at insufficient data.

  28. Q28D3 · Evaluating AI ApplicationsSelect one

    A grounded-answer eval checks whether answers are supported by retrieved context. When should it be re-run?

    • A. Whenever the retrieval pipeline, chunking, prompt or model changes
    • B. Only once, at the initial launch
    • C. Never, if the first result was good
    • D. Only when a customer complains
    Show answer

    Answer: A.

    Any change to retrieval, chunking, prompt or model can alter grounding, so the eval is re-run on each such change. B and C freeze evaluation while the system evolves. D is reactive and lets regressions ship unnoticed.

  29. Q29D4 · Designing and Building Agentic SystemsSelect one

    You want OpenAI to run and recover a durable, hours-long agent session with managed context compaction. Which runtime fits BEST?

    • A. The Agents API, OpenAI's managed harness for durable cloud sessions
    • B. The Responses API with your own loop
    • C. The open-source Agents SDK you host yourself
    • D. Chat Completions in a cron job
    Show answer

    Answer: A.

    The Agents API is the OpenAI-managed harness that runs sessions, orchestration, context compaction and recovery, which matches durable hours-long work. B and C put the durability and recovery burden on you. D has no session, recovery or compaction machinery.

  30. Q30D4 · Designing and Building Agentic SystemsSelect one

    A team wants code-first control of a custom multi-step orchestration they run on their own servers. Which runtime is the BEST fit?

    • A. The Agents SDK, an open-source framework you run yourself
    • B. The Agents API managed harness
    • C. The Batch API
    • D. Chat Completions only
    Show answer

    Answer: A.

    The Agents SDK gives code-first, self-hosted orchestration and guardrails, matching the requirement to run on their own servers. B is OpenAI-managed, not self-hosted. C is for bulk offline jobs. D lacks the orchestration and guardrail primitives the SDK provides.

  31. Q31D4 · Designing and Building Agentic SystemsSelect one

    An EU bank requires EU data residency and zero data retention. Is the Agents API eligible?

    • A. No: the Agents API is US data residency only and offers no ZDR, and a self-hosted sandbox does not change that
    • B. Yes, if they set the region header to EU
    • C. Yes, ZDR is on by default
    • D. Yes, once they enable Private Link
    Show answer

    Answer: A.

    The Agents API is US data residency only with no ZDR, and choosing a self-hosted sandbox does not lift those constraints. B invents a region header. C is false; ZDR is not offered. D confuses network privacy with residency and retention.

  32. Q32D4 · Designing and Building Agentic SystemsSelect one

    What is the correct way to let an agent read an internal repository through a standard protocol?

    • A. Expose it via an MCP server the agent connects to
    • B. Paste the whole repository into the prompt each turn
    • C. Fine-tune the model on the repository nightly
    • D. Give the agent your personal SSH key in plaintext
    Show answer

    Answer: A.

    MCP is the standard protocol for connecting tools and data sources to an agent, so an MCP server is the correct mechanism. B does not scale and leaks context budget. C is costly and stale between runs. D is an insecure credential-handling anti-pattern.

  33. Q33D4 · Designing and Building Agentic SystemsSelect one

    An agent can issue refunds. What control is REQUIRED before it acts on a high-value refund?

    • A. A human-in-the-loop approval gate before the irreversible action executes
    • B. A higher reasoning effort setting
    • C. A larger context window
    • D. A friendlier system prompt
    Show answer

    Answer: A.

    Irreversible, high-value actions require a human approval gate before execution. B, C and D may affect quality or capacity but none of them stops the agent from committing an unwanted irreversible action.

  34. Q34D4 · Designing and Building Agentic SystemsSelect two

    Which TWO statements about the Agents API managed harness are TRUE?

    • A. OpenAI runs the session, orchestration, context compaction and recovery
    • B. It supports subagent delegation via multi_agent with a bounded concurrency setting
    • C. It offers EU data residency and ZDR
    • D. You must implement session recovery yourself
    • E. It is the recommended choice when you need full self-hosted execution
    Show answer

    Answer: A and B.

    The managed harness runs the session lifecycle and supports subagent delegation via multi_agent with max_concurrent_subagents. C is false: it is US-only with no ZDR. D contradicts the managed recovery it provides. E describes the Agents SDK, not the Agents API.

  35. Q35D4 · Designing and Building Agentic SystemsSelect one

    A simple assistant needs to call one internal function and return an answer, with full control over each step and no session or sandbox. Which runtime is simplest and correct?

    • A. The Responses API with function calling, owning the loop yourself
    • B. The Agents API managed harness
    • C. The Agents SDK with orchestration and guardrails
    • D. Fine-tuning
    Show answer

    Answer: A.

    For a single tool call with full step-by-step control and no session or sandbox needs, the Responses API with function calling is the simplest fit. B and C add managed sessions or framework machinery the task does not need. D is a training operation, not a runtime for tool use.

  36. Q36D4 · Designing and Building Agentic SystemsSelect one

    What is the difference between multi-agent handoffs and subagents?

    • A. A handoff transfers control to another agent; a subagent is delegated a bounded sub-task that runs and returns to the parent
    • B. They are the same feature under two names
    • C. Subagents replace the parent permanently
    • D. Handoffs only work in the Batch API
    Show answer

    Answer: A.

    A handoff passes control to a different agent, while a subagent is delegated a bounded piece of work and returns its result to the parent. B denies a real distinction. C misdescribes delegation. D places handoffs in an unrelated API.

  37. Q37D4 · Designing and Building Agentic SystemsSelect one

    A support agent needs to read orders, but a developer grants it write access to billing 'just in case'. What principle is violated?

    • A. Least privilege: grant only the permissions the task requires
    • B. Idempotency
    • C. Statelessness
    • D. Backpressure
    Show answer

    Answer: A.

    Granting more access than the task needs violates least privilege; the fix is read-only order access. B, C and D are real engineering concepts but none of them describes over-broad permissions.

  38. Q38D4 · Designing and Building Agentic SystemsSelect one

    A hosted sandbox in the Agents API is billed how, relative to a self-hosted one?

    • A. Model API rates plus standard tool rates plus container rates for the hosted sandbox
    • B. Free, because it is managed
    • C. A flat monthly fee only
    • D. The same as a self-hosted sandbox with no container charge
    Show answer

    Answer: A.

    Agents API billing is model rates plus standard tool rates plus container rates for hosted sandboxes. B ignores the container charge. C invents a flat fee. D is wrong because the container rate is exactly what a hosted sandbox adds over self-hosting.

  39. Q39D4 · Designing and Building Agentic SystemsSelect one

    An agent occasionally takes a wrong tool action and afterwards nobody can tell why. What is missing?

    • A. Observability: tracing of the agent's steps, tool calls and inputs/outputs
    • B. A larger model
    • C. More subagents
    • D. A higher spend limit
    Show answer

    Answer: A.

    Being unable to explain a past action is an observability gap; tracing steps and tool calls provides the audit trail. B, C and D add capability, parallelism or budget but none of them records what happened for later diagnosis.

  40. Q40D5 · Retrieval-Augmented GenerationSelect one

    A knowledge base of 50,000 internal documents changes weekly and answers must cite sources. What is the BEST approach?

    • A. Retrieval-augmented generation over an indexed vector store with citations
    • B. Fine-tune the model weekly on all 50,000 documents
    • C. Put all documents in the prompt every request
    • D. Rely on the model's training knowledge
    Show answer

    Answer: A.

    RAG retrieves current documents at query time and can attach citations, handling weekly change cheaply. B is expensive, slow to update and hard to cite. C is infeasible and wasteful. D cannot know private, changing internal content.

  41. Q41D5 · Retrieval-Augmented GenerationSelect one

    Users search by exact error codes like ERR-4021 and pure semantic retrieval keeps missing them. What is the fix?

    • A. Add keyword/lexical retrieval to complement semantic search (hybrid retrieval)
    • B. Increase the embedding dimension
    • C. Switch to a bigger generation model
    • D. Raise the temperature
    Show answer

    Answer: A.

    Exact tokens like error codes are matched by lexical search, so hybrid retrieval combining keyword and semantic recall fixes the misses. B tweaks embeddings but still favours meaning over exact strings. C and D change generation, not retrieval.

  42. Q42D5 · Retrieval-Augmented GenerationSelect one

    Clinicians must only retrieve protocols for their own department. Where must this restriction be enforced?

    • A. As access-control filtering in the retrieval layer, not by prompt instructions alone
    • B. By asking the model not to reveal other departments' protocols
    • C. By raising reasoning effort
    • D. By reducing chunk size
    Show answer

    Answer: A.

    Authorisation must be enforced in retrieval so out-of-scope documents are never fetched; a prompt request is bypassable. B relies on the model behaving, which is not a security control. C and D affect quality or cost, not access.

  43. Q43D5 · Retrieval-Augmented GenerationSelect one

    A retrieved chunk lacks the answer, yet the model responds with a confident number from training knowledge. How do you prevent this?

    • A. Instruct the model to answer only from retrieved context and to say it does not know when context is insufficient, and eval for grounding
    • B. Use a bigger model
    • C. Increase max output tokens
    • D. Turn off retrieval
    Show answer

    Answer: A.

    Grounding requires instructing the model to answer only from context and to abstain otherwise, verified by a grounding eval. B may still hallucinate. C affects length, not grounding. D removes the very context that anchors the answer.

  44. Q44D5 · Retrieval-Augmented GenerationSelect one

    Chunks are set to 4,000 tokens each and answers now include lots of irrelevant text. What is the MOST likely cause?

    • A. Chunks are too large, so each retrieved chunk carries unrelated content
    • B. The embedding model is broken
    • C. The generation model is too small
    • D. Temperature is too low
    Show answer

    Answer: A.

    Oversized chunks bundle unrelated passages, so retrieval brings in noise; smaller, focused chunks help. B is unlikely to manifest as bloated-but-relevant retrieval. C and D do not explain why retrieved text is off-topic.

  45. Q45D5 · Retrieval-Augmented GenerationSelect one

    Why is overlap added between adjacent chunks when indexing?

    • A. So an answer spanning a boundary is not split and lost between two chunks
    • B. To reduce total storage
    • C. To make embeddings deterministic
    • D. To disable semantic search
    Show answer

    Answer: A.

    Overlap preserves context that straddles a chunk boundary so a boundary-spanning answer is still retrievable. B is false; overlap increases storage. C and D are unrelated to what overlap does.

  46. Q46D5 · Retrieval-Augmented GenerationSelect two

    Which TWO metrics should a grounded-answer eval measure?

    • A. Retrieval quality (did the right context get fetched)
    • B. Faithfulness (is the answer supported by the retrieved context)
    • C. Tokens per second of the embedding model
    • D. The color scheme of the UI
    • E. The size of the vector store on disk
    Show answer

    Answer: A and B.

    A grounded-answer eval measures whether retrieval fetched the right context and whether the answer is faithful to it. C is a throughput stat, not answer quality. D is unrelated. E is a storage metric, not a quality measure.

  47. Q47D5 · Retrieval-Augmented GenerationSelect one

    A developer proposes putting a whole 900-page manual in the prompt every request because the window is 1.05M tokens. Why is RAG usually better here?

    • A. RAG sends only relevant passages, cutting per-request cost and latency and reducing distraction
    • B. RAG is always more accurate on every query
    • C. The context window is actually too small for the manual
    • D. Prompts cannot contain documents
    Show answer

    Answer: A.

    Even when the manual fits, stuffing it every request pays for and processes irrelevant tokens on each call; RAG retrieves only what a query needs. B overclaims universal accuracy. C is false since 900 pages fit in 1.05M tokens. D is untrue.

  48. Q48D5 · Retrieval-Augmented GenerationSelect one

    What does file search over a managed vector store give you out of the box?

    • A. Hosted chunking, embedding, indexing and retrieval you can call as a tool
    • B. Automatic fine-tuning of the base model
    • C. A guarantee of zero hallucination
    • D. Free unlimited storage
    Show answer

    Answer: A.

    Managed file search provides hosted chunking, embedding, indexing and retrieval exposed as a tool. B is a different, training operation. C is not something any retrieval layer can guarantee. D invents a pricing claim.

  49. Q49D6 · Performance, Latency and CostSelect one

    Many requests share a 5,000-token system prompt followed by a short user message. What reduces cost and latency MOST directly?

    • A. Prompt caching of the shared prefix
    • B. Raising reasoning effort
    • C. Switching to a larger model
    • D. Increasing max output tokens
    Show answer

    Answer: A.

    A large, repeated prefix is exactly what prompt caching targets, cutting the cost and time of reprocessing it. B and C increase spend and latency. D affects output length, not the repeated input.

  50. Q50D6 · Performance, Latency and CostSelect one

    A nightly job classifies 5M documents and latency does not matter. Which processing choice is MOST cost-effective?

    • A. The Batch API
    • B. Synchronous requests on gpt-6-astra
    • C. Fast mode on every request
    • D. The Realtime API
    Show answer

    Answer: A.

    Latency-insensitive bulk work is the Batch API's purpose and it is the cheapest tier for it. B pays a premium model synchronously. C and D optimise for interactivity the job does not need, at higher cost.

  51. Q51D6 · Performance, Latency and CostSelect one

    A summariser runs 200,000 times per day at 1,000 input and 200 output tokens on gpt-5.6-luna ($0.20/M input, $1.20/M output). Roughly what is the daily cost?

    • A. About $88 per day
    • B. About $8 per day
    • C. About $440 per day
    • D. About $1,760 per day
    Show answer

    Answer: A.

    Input: 200,000 x 1,000 = 200M x $0.20/M = $40. Output: 200,000 x 200 = 40M x $1.20/M = $48. Total ~$88/day, so A. B is a tenth of the real figure, C uses Terra-scale rates, and D is far too high.

  52. Q52D6 · Performance, Latency and CostSelect one

    An interactive feature 'feels slow'. Before changing anything, what should you do FIRST?

    • A. Measure latency (including time-to-first-token) to find where the time actually goes
    • B. Immediately switch to the cheapest model
    • C. Increase reasoning effort
    • D. Add more subagents
    Show answer

    Answer: A.

    You cannot optimise what you have not measured, so profiling latency and time-to-first-token comes first. B, C and D are blind changes that may not touch the real bottleneck and could make it worse.

  53. Q53D6 · Performance, Latency and CostSelect two

    Time-to-first-token dominates latency for an interactive feature running gpt-5.6-sol on high effort. Which TWO levers most directly help?

    • A. Lower the reasoning effort to the least that still passes the eval
    • B. Stream the response so users see output as it is produced
    • C. Increase max output tokens
    • D. Move the workload to the Batch API
    • E. Add more retrieved context to every prompt
    Show answer

    Answer: A and B.

    Lower effort cuts the thinking that delays the first token, and streaming shows output as it arrives, improving perceived latency. C lengthens generation. D is for offline bulk work, not interactive latency. E adds input processing, increasing latency.

  54. Q54D6 · Performance, Latency and CostSelect one

    You are editing a large document where about 95% of the output equals the input. Which feature cuts latency MOST?

    • A. Predicted outputs
    • B. The Batch API
    • C. Higher reasoning effort
    • D. A larger context window
    Show answer

    Answer: A.

    Predicted outputs accelerate generation when most of the output is already known, exactly the edit-with-small-changes case. B is for offline bulk jobs. C adds latency. D affects capacity, not generation speed for near-identical output.

  55. Q55D6 · Performance, Latency and CostSelect one

    A latency-tolerant workload wants lower cost than standard processing but cannot wait the hours the Batch API takes. Which tier fits?

    • A. Flex processing
    • B. Fast mode
    • C. The Realtime API
    • D. Synchronous standard requests
    Show answer

    Answer: A.

    Flex processing trades some latency for lower cost without the multi-hour turnaround of Batch, fitting a latency-tolerant but not offline workload. B optimises for speed at higher cost. C targets real-time interaction. D is the standard tier it is trying to beat on cost.

  56. Q56D7 · Production Safety and OperationsSelect one

    An API call returns a 429. What is the correct handling?

    • A. Back off and retry with exponential backoff and jitter
    • B. Retry immediately in a tight loop
    • C. Treat it as a fatal error and stop the service
    • D. Switch models permanently
    Show answer

    Answer: A.

    A 429 is a rate-limit signal handled by exponential backoff with jitter. B worsens the overload. C overreacts to a transient condition. D does not address rate limiting and may just move the problem.

  57. Q57D7 · Production Safety and OperationsSelect one

    Which error should you NOT automatically retry?

    • A. A 400 invalid-request error caused by a malformed payload
    • B. A 429 rate-limit error
    • C. A 500 transient server error
    • D. A 503 service-unavailable error
    Show answer

    Answer: A.

    A 400 means the request itself is wrong, so retrying the same payload will fail again; fix the request instead. B, C and D are transient or throttling conditions where backoff and retry are appropriate.

  58. Q58D7 · Production Safety and OperationsSelect one

    A team fears a runaway agent loop could generate a huge bill. Which control MOST directly caps the dollar cost?

    • A. Spend limits on the account or project
    • B. A friendlier system prompt
    • C. A larger context window
    • D. Higher reasoning effort
    Show answer

    Answer: A.

    Spend limits directly cap dollar exposure regardless of loop behaviour. B, C and D have no bearing on total spend and D would increase it. Rate limits help throttle but spend limits cap the money most directly.

  59. Q59D7 · Production Safety and OperationsSelect two

    A Kubernetes workload should call the API without holding a long-lived API key. Which TWO practices are correct?

    • A. Use workload identity federation so the pod exchanges its identity for short-lived credentials
    • B. Scope the credentials to least privilege for the job
    • C. Bake a static API key into the container image
    • D. Store one shared key in the manifest for all clusters
    • E. Disable logging so the key cannot be traced
    Show answer

    Answer: A and B.

    Workload identity federation avoids long-lived keys by issuing short-lived credentials, and least-privilege scoping limits what those credentials can do. C embeds a secret in an image that can leak. D widens blast radius and complicates rotation. E removes observability and does nothing about the key itself.

  60. Q60D7 · Production Safety and OperationsSelect two

    An agent processes untrusted customer messages and can take actions. Which TWO controls address prompt-injection risk BEST?

    • A. Least-privilege tool scopes so a hijacked agent can do little
    • B. Human approval gates before irreversible actions
    • C. Raising the temperature
    • D. Using a larger model
    • E. Increasing max output tokens
    Show answer

    Answer: A and B.

    Least privilege limits the blast radius of a successful injection and approval gates stop irreversible actions from firing automatically. C changes randomness, D does not prevent injection, and E only affects output length.

Last updated Sep 18, 2026