AI Cert Prep
Type to search documentation.

Codex Path

Codex · Mock Exam 2

A harder 50-item, domain-weighted independent mock exam for the OpenAI Academy Codex pathway, with full explanations and a readiness readout. Not an official OpenAI assessment.

This is the second full-length, domain-weighted independent mock exam for the Codex track. It is built from publicly available OpenAI learning objectives and is not an official OpenAI assessment. Mock Exam 2 is deliberately harder than Mock Exam 1: more multi-constraint stems and more FIRST / BEST / MOST cost-effective / TWO qualifiers, with more scenario framing. All 50 questions are new and distinct from Mock Exam 1 and the domain-page items. Use it timed as your go/no-go gate.

Instructions

  • Time: 60 minutes (our design choice for a 50-item mock; the Academy assessments are shorter, randomised selections from a larger bank).
  • Items: 50, multiple-choice and multiple-response. Each item states how many answers to select.
  • Selection: for multiple-response items you must select all correct options and no incorrect ones; partial selections are marked wrong.
  • No guessing penalty: answer every question.
  • Target: aim for at least 80% raw (≈ 40/50) before taking the real Academy Codex assessment, which passes at ≥ 80%.
  • Work each question before expanding the answer.

Domain distribution

#DomainItems here
1Codex Fundamentals and Surfaces10
2Core Coding Workflows12
3Extending and Configuring Codex10
4Team Adoption and Governance10
5Scaling Across Teams and Systems8

Total: 10 + 12 + 10 + 10 + 8 = 50 items.

Readiness interpretation

This is an independent readiness indicator, not a score and not a prediction of any official result.

Raw score (of 50)BandInterpretation
45–50Strong readiness90%+; strong across all domains
40–44Assessment ready80–89%; at or above the Academy threshold
35–39Building confidence70–79%; close, target your weak domains
under 35Keep learningbelow 70%; revisit the domain pages before re-attempting

The 80% line is deliberate: it matches the Academy badge threshold.

Take the mock exam

Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · Codex Fundamentals and SurfacesSelect one

    A one-file typo fix must ship in the next five minutes, cost is a concern, and it decomposes into no sub-tasks. Which model and reasoning effort are MOST cost-effective?

    • A. gpt-6-astra at Ultra
    • B. gpt-5.6-luna at Low or Medium
    • C. gpt-5.6-sol at Max
    • D. gpt-5.6-terra at Extra High
    Show answer

    Answer: B.

    A trivial mechanical fix wants the cheapest model (Luna) at low effort. Astra at Ultra (A) burns budget and adds pointless parallelism, Sol at Max (C) over-thinks a typo, and Terra at Extra High (D) still applies far more effort than a typo warrants.

  2. Q2D1 · Codex Fundamentals and SurfacesSelect one

    An engineer must decide between Max and Ultra for a job that is a single, indivisible constraint-solving problem that keeps timing out at High. Which is the BEST choice and why?

    • A. Ultra, because it is the highest rung and always strongest
    • B. Max, because the problem is one task needing more thinking time, not independent parts to parallelise
    • C. Ultra, because a single hard problem is exactly what subagents are for
    • D. Max, because it uses a larger model than Ultra
    Show answer

    Answer: B.

    Max deepens reasoning on a single indivisible task, which is this case. Ultra (A, C) delegates to parallel subagents and does not help an indivisible problem, and neither Max nor Ultra changes the model choice (D).

  3. Q3D1 · Codex Fundamentals and SurfacesSelect one

    A team's shared config pins gpt-5.4, everyone signs in with ChatGPT, and after 31 August 2026 the model stops resolving. Which single action FIRST restores work with the direct replacement?

    • A. Renew the ChatGPT subscription
    • B. Set model to gpt-5.6-terra in the shared config.toml
    • C. Regenerate the corrupted config file
    • D. Upgrade the Codex CLI to the latest version
    Show answer

    Answer: B.

    gpt-5.4 retired under ChatGPT sign-in and Terra is its direct replacement, so updating the shared default resolves it. A subscription renewal (A), a corrupt file (C) and a CLI upgrade (D) do not explain a model-specific cut-off tied to that date and sign-in type.

  4. Q4D1 · Codex Fundamentals and SurfacesSelect two

    A workload is high-volume, mechanical, latency-sensitive and cost-constrained, running unattended in CI on GitHub. Which TWO choices best fit it?

    • A. The Codex CLI invoked with codex exec
    • B. gpt-5.6-luna as the model
    • C. The IDE extension in interactive mode
    • D. gpt-6-astra as the model
    • E. Ultra reasoning by default
    Show answer

    Answer: A and B.

    Unattended CI is the codex exec case (A) and a high-volume mechanical job wants the cheapest model, Luna (B). The interactive IDE (C) waits for input, Astra (D) overspends, and Ultra (E) adds parallelism that mechanical volume does not need.

  5. Q5D1 · Codex Fundamentals and SurfacesSelect one

    A user cannot find Ultra in their model picker but their task genuinely decomposes into independent parts. What is the correct FIRST step before redesigning the task?

    • A. Reinstall the Codex CLI
    • B. Enable the 'Ultra in model picker slider' via Settings then Configuration
    • C. Upgrade to GPT-6 Astra to unlock Ultra
    • D. Move the work to Codex cloud, the only place Ultra exists
    Show answer

    Answer: B.

    When Ultra is absent from the picker it is enabled via Settings then Configuration then the 'Ultra in model picker slider'. Reinstalling (A) and upgrading a model (C) are unrelated, and Ultra is not restricted to cloud (D).

  6. Q6D1 · Codex Fundamentals and SurfacesSelect one

    Which statement most accurately distinguishes GPT-5.6 Sol from GPT-5.6 Terra for Codex work?

    • A. Sol is the cheapest high-volume model; Terra is the most expensive
    • B. Sol is for complex, open-ended, high-value work; Terra is the pragmatic all-rounder and natural GPT-5.5 replacement
    • C. Sol is text-only research preview; Terra is a vision model
    • D. They are aliases for the same model
    Show answer

    Answer: B.

    Sol targets complex, open-ended, high-value work while Terra is the pragmatic default and the natural GPT-5.5 replacement. The cheapest high-volume model is Luna not Sol (A), the text-only preview is Spark (C), and they are distinct models (D).

  7. Q7D1 · Codex Fundamentals and SurfacesSelect one

    A lead insists that because the desktop app and CLI 'look and feel different', each must be a different underlying agent needing separate configuration. What is the MOST accurate correction?

    • A. They are right; treat them as separate agents
    • B. They are one agent behind different surfaces sharing a single config.toml; only the interaction model differs
    • C. Only the CLI is a real agent; the desktop app is a viewer
    • D. The desktop app cannot run Codex at all
    Show answer

    Answer: B.

    Codex is one agent behind many surfaces sharing one config.toml; the surfaces differ only in interaction. They are not separate agents (A), the desktop app is a real Codex surface not a viewer (C), and it does run Codex (D).

  8. Q8D1 · Codex Fundamentals and SurfacesSelect one

    Three tasks are queued: (1) an overnight mechanical rename across many files, (2) one hard indivisible algorithm redesign, (3) four unrelated independent bug fixes due today. Which allocation of reasoning effort is BEST?

    • A. All three at Ultra to maximise strength
    • B. Rename at Low/Medium, redesign at Max, the four fixes with Ultra
    • C. Rename at Ultra, redesign at Low, fixes at Medium
    • D. All three at Max
    Show answer

    Answer: B.

    The mechanical rename needs little effort, the single hard redesign wants Max's deeper single-task thinking, and the four independent fixes decompose so they fit Ultra's parallel subagents. Ultra everywhere (A) and Max everywhere (D) misapply effort, and C inverts the correct assignments.

  9. Q9D1 · Codex Fundamentals and SurfacesSelect one

    Which surface should a developer reach for when they want the diff in front of them in their editor and to approve each change during a hard, interactive refactor?

    • A. Codex cloud
    • B. The Codex IDE extension
    • C. codex exec on the CLI
    • D. Codex Micro
    Show answer

    Answer: B.

    The IDE extension puts the diff in front of you for a tight, interactive edit-review loop. Cloud (A) suits long-running or parallel work away from the machine, codex exec (C) is non-interactive, and Micro (D) is for small quick tasks.

  10. Q10D1 · Codex Fundamentals and SurfacesSelect one

    Which is the correct way to select the model and effort separately in Codex?

    • A. Model with the Low-to-Ultra ladder; effort with -m or /model
    • B. Model with -m, --model or /model; effort with the Low-to-Ultra reasoning ladder
    • C. Both model and effort are set only in AGENTS.md
    • D. Both are set with the single --effort flag
    Show answer

    Answer: B.

    Model is chosen with -m, --model or /model, while effort is the Low-to-Ultra ladder — two independent controls. Option A swaps the two, AGENTS.md is repo guidance not the effort control (C), and there is no single --effort flag for both (D).

  11. Q11D2 · Core Coding WorkflowsSelect one

    Codex returns a 300-line diff touching four unrelated areas from the one-line prompt 'speed up the export', and claims a 40% improvement with green tests, with a demo in one hour. What should you do FIRST?

    • A. Merge it; green tests and a deadline justify shipping
    • B. Re-scope the task with a specific goal, boundaries and definition of done, then verify independently
    • C. Switch to Astra and re-run the same one-liner
    • D. Cherry-pick the parts of the diff that look safe and merge those
    Show answer

    Answer: B.

    A sprawling diff from a vague prompt is a scoping failure; re-scoping and independent verification is the correct first move, and the deadline raises rather than lowers the cost of a bad merge. Merging on the claim (A) skips verification, a bigger model (C) still sprawls on a vague task, and cherry-picking (D) risks an inconsistent change.

  12. Q12D2 · Core Coding WorkflowsSelect two

    A reviewer wants to trust a Codex PR in minutes without reconstructing the author's confidence. Which TWO pieces of evidence are MOST valuable?

    • A. A focused diff with any unexpected changes explicitly called out
    • B. The reasoning effort level used
    • C. A passing test run or CI link plus a regression test pinning the new behaviour
    • D. A statement that Codex is a capable model
    • E. The number of tokens the run consumed
    Show answer

    Answer: A and C.

    A focused diff (A) and a passing test run with a regression test (C) let a reviewer judge scope and correctness quickly. The effort level (B), a claim about the model (D) and token counts (E) tell a reviewer nothing about whether the change is right.

  13. Q13D2 · Core Coding WorkflowsSelect one

    A nightly maintenance job must apply a mechanical migration across a repo with no human present, on the lowest cost that works, and it must not be able to escalate silently. Which setup is BEST?

    • A. Interactive codex at High effort so someone can approve escalations later
    • B. codex exec with gpt-5.6-luna, a restrictive permission mode and a sandbox
    • C. The IDE extension left open overnight with broad auto-approve
    • D. codex exec with gpt-6-astra and broad auto-approve for speed
    Show answer

    Answer: B.

    Unattended mechanical work is the codex exec case with a cheap model (Luna) and restrictive permissions plus a sandbox so it cannot escalate silently. Interactive mode (A) waits for input, an open IDE (C) is not for unattended runs and broad auto-approve is unsafe, and Astra with broad auto-approve (D) both overspends and removes containment.

  14. Q14D2 · Core Coding WorkflowsSelect one

    A teammate argues that because a Codex change has green tests, reading the diff is unnecessary before merge. What is the strongest counter-argument?

    • A. Green tests are a full review, so the teammate is correct
    • B. Tests are one evidence type; the diff must be read for scope creep and to confirm the change matches the intended, bounded task
    • C. Tests are unreliable and should be ignored
    • D. Only the model ID matters for review
    Show answer

    Answer: B.

    Tests are necessary but not sufficient; the diff reveals scope creep and confirms the change is on-task. Green tests are not a full review (A), tests remain valuable (C), and the model ID is irrelevant to correctness (D).

  15. Q15D2 · Core Coding WorkflowsSelect one

    Codex reports a bug is fixed and the suite is green, but the original reported repro was never added to the suite. What is the MOST likely hidden risk?

    • A. No risk; a green suite proves the specific fix
    • B. Codex may have fixed a lookalike issue while the real repro still fails; reproduce it before and after and add a regression test
    • C. The suite is too slow and should have tests deleted
    • D. The wrong model was used; switch models
    Show answer

    Answer: B.

    A green suite that never exercised the real repro can hide a lookalike fix, so reproduce the reported case and add a regression test. A green suite alone does not prove the specific fix (A), deleting tests (C) reduces coverage, and swapping models (D) does not verify anything.

  16. Q16D2 · Core Coding WorkflowsSelect one

    Three engineers must build three independent features locally at once without their branches colliding, keeping each verifiable. Which approach is BEST?

    • A. Work through the features one at a time on the same branch
    • B. Use a separate git worktree or branch per feature (or parallel cloud tasks), verifying each before a controlled merge
    • C. Commit all three directly to main as they finish
    • D. Disable the tests so all three merge faster
    Show answer

    Answer: B.

    Separate worktrees or branches (or parallel cloud tasks) let independent work proceed at once, verified before a controlled merge. Serialising (A) wastes the parallelism, committing to main (C) is unsafe, and disabling tests (D) removes verification.

  17. Q17D2 · Core Coding WorkflowsSelect one

    Which practice most directly satisfies the Codex objective to 'verify changes and provide clear evidence for review'?

    • A. Choosing the most expensive model available
    • B. Running the tests yourself, reading the diff, and attaching the passing run plus a regression test to the PR
    • C. Writing the longest possible prompt
    • D. Merging quickly to keep momentum
    Show answer

    Answer: B.

    Independently verifying and attaching evidence is exactly what the objective asks for. Model choice (A) and prompt length (C) do not produce evidence, and merging quickly (D) skips verification.

  18. Q18D2 · Core Coding WorkflowsSelect one

    A prompt says 'improve the auth code'. Which rewrite BEST turns it into a scoped, reviewable task?

    • A. 'Make the auth code much better and faster overall'
    • B. 'Fix the token-refresh race in auth/refresh.ts; do not touch the public API or DB schema; the new test refresh.race.test.ts must pass and existing tests must stay green'
    • C. 'Refactor everything under auth/ however you see fit'
    • D. 'Use the strongest model on the auth code'
    Show answer

    Answer: B.

    Option B supplies a goal, boundaries and an executable definition of done. 'Better and faster overall' (A) and 'refactor everything however you see fit' (C) are still unscoped, and specifying a model (D) is not scoping the task.

  19. Q19D2 · Core Coding WorkflowsSelect one

    A developer wants runtime signals — test output and stack traces — fed back into the loop to verify and debug a change. Which context source provides that BEST?

    • A. The PR description
    • B. The integrated terminal
    • C. A comment in the source file
    • D. The workspace analytics dashboard
    Show answer

    Answer: B.

    The integrated terminal surfaces runtime signals such as test output and stack traces for verifying and debugging in-loop. A PR description (A) and a source comment (C) are static text, and the analytics dashboard (D) reports adoption, not runtime signals.

  20. Q20D2 · Core Coding WorkflowsSelect one

    For a bug reported as 'checkout occasionally double-charges', which is the MOST reliable definition of done to hand Codex?

    • A. 'Stop the double-charge somehow'
    • B. A failing test that reproduces the double-charge, required to pass while the full suite stays green, plus a note not to change the payment gateway contract
    • C. 'Rewrite the checkout module from scratch'
    • D. 'Add logging and hope the problem goes away'
    Show answer

    Answer: B.

    A failing test that reproduces the reported behaviour is an executable definition of done, and the boundary protects the gateway contract. 'Somehow' (A) is unscoped, a full rewrite (C) is disproportionate scope creep, and adding logging (D) does not define done or fix the bug.

  21. Q21D2 · Core Coding WorkflowsSelect one

    A developer habitually runs CI fix jobs by opening an interactive codex session on their laptop and typing the task. Why is this the wrong tool, and what is the fix?

    • A. It is fine; interactive sessions are ideal for CI
    • B. CI is unattended and scripted, so it should use non-interactive codex exec rather than an interactive session that waits for input
    • C. The laptop is too slow; buy a faster one
    • D. Interactive sessions cannot run tests
    Show answer

    Answer: B.

    CI runs without a human, so codex exec is the correct non-interactive form; an interactive session would block waiting for input. Interactive is not ideal for CI (A), the issue is the mode not hardware (C), and interactive sessions can run tests (D).

  22. Q22D2 · Core Coding WorkflowsSelect two

    Which TWO elements are core parts of a well-scoped Codex task?

    • A. Explicit boundaries on what must not be touched
    • B. The most expensive available model
    • C. A definition of done, such as a specific test that must pass
    • D. The longest possible prompt
    • E. Ultra reasoning by default
    Show answer

    Answer: A and C.

    Scope is goal plus boundaries plus definition of done plus constraints; boundaries (A) and a definition of done (C) are two of those. A pricier model (B), a longer prompt (D) and defaulting to Ultra (E) are not scoping and do not substitute for it.

  23. Q23D3 · Extending and Configuring CodexSelect one

    A fintech team wants a nightly dependency-upgrade job that opens a PR, and proposes broad auto-approve, full cloud internet access, experimental context management, and auto-merge on auto-review — on a ChatGPT Enterprise workspace. Which single element is impossible on availability grounds alone?

    • A. Broad auto-approve
    • B. Full cloud internet access
    • C. Experimental context management on an Enterprise workspace
    • D. Auto-merge on auto-review
    Show answer

    Answer: C.

    Experimental context management is Plus/Pro sign-in only at launch, so it is unavailable on an Enterprise workspace regardless of its merits. Broad auto-approve (A), full internet access (B) and auto-merge on auto-review (D) are unsafe choices but are not blocked purely by availability.

  24. Q24D3 · Extending and Configuring CodexSelect two

    For a risky unattended run, which TWO controls should be combined to limit both what happens without a human and how far any action can reach?

    • A. A restrictive permission mode
    • B. Broad auto-approve
    • C. A sandbox that contains the blast radius
    • D. Full cloud internet access
    • E. Disabling all approvals
    Show answer

    Answer: A and C.

    A restrictive permission mode (A) governs what needs approval while a sandbox (C) contains impact; together they limit the decision surface and the blast radius. Broad auto-approve (B), full internet access (D) and disabling approvals (E) increase risk rather than contain it.

  25. Q25D3 · Extending and Configuring CodexSelect one

    A team needs Codex to reach an internal data warehouse, run a custom check before every write, AND reproduce a specific session for debugging. Which mapping of needs to extension points is correct?

    • A. Data warehouse: hook; pre-write check: MCP; reproduce session: skill
    • B. Data warehouse: MCP; pre-write check: hook; reproduce session: record and replay
    • C. All three: a single plugin
    • D. Data warehouse: record and replay; pre-write check: skill; reproduce session: MCP
    Show answer

    Answer: B.

    External data is MCP, running your own logic at a lifecycle point is a hook, and reproducing a session is record and replay. Option A swaps MCP and hooks, a single plugin (C) does not map to these distinct needs, and D mismatches every mapping.

  26. Q26D3 · Extending and Configuring CodexSelect one

    An engineer proposes relying on auto-review as the sole gate before a fintech dependency PR auto-merges. What is the MOST accurate objection?

    • A. None; auto-review is a complete substitute for a human reviewer
    • B. Auto-review produces useful evidence but removing the human gate is unsafe; keep a required human review before merge
    • C. Auto-review only checks formatting, so use a linter instead
    • D. Auto-review is unavailable in the cloud
    Show answer

    Answer: B.

    Auto-review is evidence feeding a human gate, not a replacement for it, especially for a regulated change. It does not fully substitute for a reviewer (A), it is not limited to formatting (C), and it is available across Codex surfaces including cloud (D).

  27. Q27D3 · Extending and Configuring CodexSelect one

    What is the correct way to enable experimental context management, and under what limit?

    • A. It is on automatically for Astra with no limits
    • B. Opt in via features.context_management.experimental_mode = true in config.toml, subject to the Plus/Pro sign-in limit at launch
    • C. Pass --context on every command; no sign-in limit
    • D. Enable it in AGENTS.md; Enterprise only
    Show answer

    Answer: B.

    It is opt-in through features.context_management.experimental_mode = true and, at launch, limited to Plus/Pro sign-in. It is not automatic (A), not a per-command flag (C), and not enabled in AGENTS.md nor Enterprise-scoped (D).

  28. Q28D3 · Extending and Configuring CodexSelect one

    A team keeps putting per-repository build commands into config.toml and wonders why behaviour differs across repos. Which correction is MOST accurate?

    • A. config.toml is per-repo, so each repo just needs its own copy
    • B. Per-repo build commands and conventions belong in each repo's AGENTS.md; config.toml sets machine or workspace-wide defaults
    • C. Build commands cannot be configured
    • D. Both files must contain identical content to work
    Show answer

    Answer: B.

    Repo-specific build commands and conventions go in AGENTS.md, while config.toml sets machine or workspace defaults, which explains the divergence. config.toml is not per-repo (A), build commands are configurable (C), and the files serve different purposes so need not match (D).

  29. Q29D3 · Extending and Configuring CodexSelect one

    A cloud dependency-fetch task needs network access to one package registry. What is the MOST appropriate internet-access configuration?

    • A. Leave internet access fully open for the whole workspace
    • B. Grant the minimum access the task needs, deliberately, and keep the default restricted otherwise
    • C. Disable all network access, even though the fetch needs it
    • D. Switch to API-key sign-in to bypass access controls
    Show answer

    Answer: B.

    Least privilege means granting only the access the task needs and keeping the default restricted. Fully open (A) is over-permissive, blocking access the task requires (C) breaks the task, and switching sign-in to bypass controls (D) is a governance failure.

  30. Q30D3 · Extending and Configuring CodexSelect one

    An engineer wants to reproduce and share an exact Codex session with a colleague for debugging. Which feature is designed for that?

    • A. Skills
    • B. Record and replay
    • C. MCP
    • D. Auto-review
    Show answer

    Answer: B.

    Record and replay captures a session so it can be re-run or shared. Skills (A) are reusable abilities, MCP (C) provides external access, and auto-review (D) summarises a diff.

  31. Q31D3 · Extending and Configuring CodexSelect one

    A team confuses hooks and rules. Which distinction is MOST accurate?

    • A. They are the same mechanism with different names
    • B. Hooks run your own logic at lifecycle points; rules are behavioural constraints that steer or forbid agent actions
    • C. Rules reach external systems; hooks do not exist
    • D. Hooks only work in the cloud; rules only in the CLI
    Show answer

    Answer: B.

    Hooks execute your own logic at lifecycle points, while rules constrain what the agent may do. They are not the same (A), external access is MCP not rules (C), and neither is surface-restricted as described (D).

  32. Q32D3 · Extending and Configuring CodexSelect one

    Which single statement correctly separates the scope of config.toml from AGENTS.md?

    • A. config.toml is per-repo guidance; AGENTS.md is machine-wide defaults
    • B. config.toml sets machine or workspace-wide defaults (model, features, permissions); AGENTS.md carries per-repository guidance (build/test commands, conventions, no-go areas)
    • C. Both are per-repository and interchangeable
    • D. config.toml is only for the cloud surface; AGENTS.md only for the CLI
    Show answer

    Answer: B.

    config.toml sets machine or workspace defaults while AGENTS.md holds per-repository guidance — the reverse of option A. They are not interchangeable (C), and neither is restricted to a single surface (D).

  33. Q33D4 · Team Adoption and GovernanceSelect one

    A healthcare company rolling Codex out to 120 engineers has repos touching PHI and a nightly cloud job currently authenticated with one engineer's personal access token. Which single change should be the FIRST governance fix for that job?

    • A. Give the engineer admin rights so the token has more scope
    • B. Re-authenticate the org automation with workload identity federation or a service account instead of a personal token
    • C. Share the personal token with the whole team so it survives
    • D. Move the job to an interactive session
    Show answer

    Answer: B.

    A personal token ties an org automation to one person and breaks when they leave; workload identity or a service account is the correct org identity. More admin scope (A) worsens least-privilege, sharing the token (C) is a security failure, and interactive mode (D) defeats an unattended job.

  34. Q34D4 · Team Adoption and GovernanceSelect two

    A regulated PHI workload needs both correct configuration and a provable action history. Which TWO belong in the exam-correct answer?

    • A. HIPAA configuration for the PHI-handling repositories
    • B. Broad auto-approve to speed the pipeline
    • C. The Compliance API and audit events for the action record
    • D. Disabling analytics for privacy
    • E. A single shared admin account for everyone
    Show answer

    Answer: A and C.

    A regulated workload needs the right configuration (HIPAA) and an audit trail (Compliance API and audit events). Broad auto-approve (B) increases risk, disabling analytics (D) removes useful oversight, and a shared admin account (E) destroys accountability.

  35. Q35D4 · Team Adoption and GovernanceSelect one

    An organisation wants to prevent Codex configuration drift across teams while enforcing which models may be used. Which pair of controls FIRST addresses both?

    • A. Personal access tokens and record and replay
    • B. Managed configuration together with workspace model availability
    • C. The reasoning ladder and codex exec
    • D. AGENTS.md in one repo and the Slack integration
    Show answer

    Answer: B.

    Managed configuration enforces consistent settings and workspace model availability restricts which models the workspace may use. PATs and record and replay (A), the reasoning ladder and codex exec (C), and one repo's AGENTS.md with Slack (D) do not enforce org-wide config or model policy.

  36. Q36D4 · Team Adoption and GovernanceSelect one

    Which statement about Codex Security is MOST accurate for governing code produced across a team?

    • A. It is CLI-only and must be run manually after merge
    • B. It spans the plugin, CLI and cloud, offers scans, deep scans, workbench triage, fixes and hardening, and integrates with CI or GitLab CI to scan before merge
    • C. It is a separate product unrelated to Codex
    • D. It only produces vulnerability reports and cannot triage or fix
    Show answer

    Answer: B.

    Codex Security spans the plugin, CLI and cloud, with scans, triage, fixes and hardening, and integrates into CI or GitLab CI so code is scanned before merge. It is not CLI-only or post-merge-manual (A), it is part of the Codex security surface (C), and it does more than reporting (D).

  37. Q37D4 · Team Adoption and GovernanceSelect one

    Leadership asks for repeatable monthly adoption numbers for a board update. Which approach is MOST appropriate?

    • A. Ask team leads for their impressions each month
    • B. Pull the figures programmatically via the Analytics API, complemented by the workspace analytics dashboard
    • C. Estimate from the number of licenses purchased
    • D. Read a sample of engineers' shell histories
    Show answer

    Answer: B.

    The Analytics API gives repeatable programmatic figures and the dashboard complements it. Impressions (A), license-count estimates (C) and shell-history sampling (D) are unreliable and not repeatable adoption measures.

  38. Q38D4 · Team Adoption and GovernanceSelect one

    An admin argues that giving everyone admin avoids permission tickets and that PHI repos need no special handling because access is internal. Which correction addresses BOTH errors?

    • A. Both claims are fine as stated
    • B. Apply least privilege so permissions match responsibility, and apply HIPAA configuration to PHI repos regardless of internal framing
    • C. Give everyone admin but disable analytics for privacy
    • D. Keep universal admin and rely on the model to protect PHI
    Show answer

    Answer: B.

    Least privilege fixes the universal-admin error and HIPAA configuration is required for PHI regardless of internal framing. Both claims are not fine (A), disabling analytics (C) removes oversight without fixing either error, and relying on the model to protect PHI (D) is not a compliance control.

  39. Q39D4 · Team Adoption and GovernanceSelect one

    Which authentication choice is correct for each actor: a GitHub Actions pipeline, a shared org automation bot, and one engineer's personal script?

    • A. Pipeline: PAT; bot: PAT; script: workload identity
    • B. Pipeline: workload identity federation; bot: service account; script: personal access token
    • C. Pipeline: service account; bot: PAT; script: workload identity
    • D. All three: the admin's personal account
    Show answer

    Answer: B.

    A CI/cloud pipeline uses workload identity federation (no long-lived secret), a shared org automation uses a service account, and an individual's script uses a PAT. Option A and C mismatch the actors, and using the admin's personal account for everything (D) destroys accountability.

  40. Q40D4 · Team Adoption and GovernanceSelect two

    During a rollout, which TWO controls together ensure access is provisioned consistently and revoked promptly when engineers leave?

    • A. Automated groups and provisioning managing the user lifecycle
    • B. The model picker
    • C. Least-privilege roles and workspace permissions matched to responsibility
    • D. Codex Security deep scans
    • E. The Analytics API
    Show answer

    Answer: A and C.

    Automated groups and provisioning manage membership so leavers lose access (A), and least-privilege roles and permissions govern what each member may do (C). The model picker (B) selects models, Codex Security (D) scans code, and the Analytics API (E) reports adoption.

  41. Q41D4 · Team Adoption and GovernanceSelect one

    An admin wants Codex-produced code scanned before merge on every repo, with an auditable record of who did what. Which pairing is correct?

    • A. Manual scans plus the Analytics API
    • B. Codex Security in CI or GitLab CI plus the Compliance API and audit events
    • C. Record and replay plus workspace model availability
    • D. Auto-review as the only control
    Show answer

    Answer: B.

    Codex Security in CI scans before merge and the Compliance API with audit events provides the who-did-what record. Manual scans and the Analytics API (A) neither scan reliably nor audit, record and replay with model availability (C) do neither job, and auto-review alone (D) is evidence, not a scanning or audit control.

  42. Q42D4 · Team Adoption and GovernanceSelect one

    Which is the correct characterisation of managed configuration versus a shared setup document?

    • A. They achieve the same outcome
    • B. Managed configuration centrally sets and enforces settings, while a document only describes them and relies on engineers matching by hand
    • C. A document enforces settings; managed configuration only describes them
    • D. Neither can control extensions
    Show answer

    Answer: B.

    Managed configuration enforces settings centrally, whereas a document merely describes them and depends on manual compliance. They are not equivalent (A), the roles are not reversed (C), and managed configuration can control extensions (D).

  43. Q43D5 · Scaling Across Teams and SystemsSelect one

    A 40-engineer, 25-repo group reports tasks-per-week tripled while the PR revert rate quietly climbed. Which signal should drive the go/no-go decision, and what is the FIRST structural fix?

    • A. Tasks-per-week; keep scaling because activity proves value
    • B. The revert rate; reinstate the review gate, scope smaller, and measure quality alongside adoption
    • C. Tasks-per-week; upgrade everyone to Astra
    • D. The revert rate; ignore it because reverts are normal
    Show answer

    Answer: B.

    The rising revert rate is the quality signal that matters, and restoring the review gate while measuring quality is the structural fix. Tasks-per-week is activity not value (A, C), and a climbing revert trend should not be ignored (D).

  44. Q44D5 · Scaling Across Teams and SystemsSelect two

    Which TWO failure modes emerge specifically when scaling Codex across many repositories rather than in single-task use?

    • A. Configuration drift causing inconsistent agent behaviour
    • B. A single engineer writing one clear prompt
    • C. Unreviewed change volume causing the review gate to lapse
    • D. Choosing the correct model for one task
    • E. Adding an AGENTS.md to one repository
    Show answer

    Answer: A and C.

    Config drift across repos (A) and review lapsing under change volume (C) are classic at-scale failure modes. A clear single prompt (B), a correct single-task model choice (D) and adding one AGENTS.md (E) are healthy single-task practices, not scaling failures.

  45. Q45D5 · Scaling Across Teams and SystemsSelect one

    A platform team must map three needs to the right Codex integration: Codex steps in GitHub CI, embedding Codex in an internal developer portal, and acting on Linear issues. Which mapping is correct?

    • A. GitHub CI: SDK; portal: GitHub Action; Linear: Slack
    • B. GitHub CI: GitHub Action; portal: Codex SDK; Linear: the Linear integration
    • C. GitHub CI: Linear integration; portal: GitHub Action; Linear: SDK
    • D. All three: the Slack integration
    Show answer

    Answer: B.

    The GitHub Action embeds Codex in GitHub CI, the SDK builds Codex into your own tooling like a portal, and the Linear integration handles issue-driven work. Option A swaps the SDK and Action, C mismatches all three, and Slack alone (D) does none of these jobs.

  46. Q46D5 · Scaling Across Teams and SystemsSelect one

    A team on GitLab plans a business-critical launch next week relying on the Codex GitLab integration. What is the MOST responsible planning stance?

    • A. Treat it as generally available and identical to GitHub's integration
    • B. Account for beta risk in the plan, since the GitLab integration is in beta
    • C. Assume GitLab is entirely unsupported and abandon the plan
    • D. Migrate the whole org to GitHub before the launch
    Show answer

    Answer: B.

    The GitLab integration is beta, so a critical launch plan should account for beta risk. It is not GA-equivalent to GitHub (A), GitLab is supported in beta rather than unsupported (C), and a full migration (D) is disproportionate.

  47. Q47D5 · Scaling Across Teams and SystemsSelect one

    When scaling Codex, which combination BEST captures whether the rollout is actually helping?

    • A. Adoption metrics alone, since more usage means more value
    • B. Adoption metrics (active users, tasks run) paired with quality metrics (review pass rate, defects, revert rate, security findings)
    • C. Quality metrics alone, ignoring usage entirely
    • D. The reasoning-effort levels chosen by engineers
    Show answer

    Answer: B.

    Pairing adoption with quality shows whether rising usage is producing good outcomes. Adoption alone can hide falling quality (A), quality alone ignores whether the tool is used (C), and effort levels (D) do not measure rollout value.

  48. Q48D5 · Scaling Across Teams and SystemsSelect one

    A manager proposes to 'keep scaling and fix problems reactively as they appear'. Why is this the wrong stance for known scaling failure modes?

    • A. It is correct; reactive fixes are efficient at scale
    • B. At scale, reactive fixes lag the damage; collisions, drift, unreviewed volume and activity-as-value need structural mitigation (isolate, standardise, safe-merge, measure)
    • C. Because a larger model would prevent all scaling problems
    • D. Because scaling problems never actually occur
    Show answer

    Answer: B.

    Known scaling failure modes need structural mitigations because reactive fixes trail the damage they cause. Reactive-only is not efficient here (A), a larger model does not prevent process failures (C), and these problems do occur at scale (D).

  49. Q49D5 · Scaling Across Teams and SystemsSelect one

    To make parallel outputs from many repositories safe to integrate, which upstream practice matters MOST?

    • A. Letting each repo keep its own conventions for local ownership
    • B. Standardising a shared AGENTS.md template and managed configuration so each output was produced under the same rules, with drift checks
    • C. Using the largest model in every repo
    • D. Merging all repos into a single monorepo overnight
    Show answer

    Answer: B.

    Standardising AGENTS.md and configuration means outputs are produced under the same rules, which is what makes parallel integration safe. Divergent conventions (A) make integration unsafe, a bigger model (C) does not fix inconsistent context, and a rushed monorepo merge (D) is disproportionate and risky.

  50. Q50D5 · Scaling Across Teams and SystemsSelect two

    Security debt accumulates when vulnerabilities merge faster than they are found across 30 repositories. Which TWO practices form the correct structural mitigation?

    • A. Run Codex Security in CI on every repository so code is scanned before it merges
    • B. Scan only the largest repository occasionally
    • C. Keep a human review gate rather than skipping review under volume
    • D. Wait for a breach, then scan everything
    • E. Trust the model and skip scanning to keep velocity
    Show answer

    Answer: A and C.

    Scanning every repo in CI before merge (A) and keeping the human review gate under volume (C) are the structural mitigations. Scanning one repo occasionally (B) leaves gaps, post-breach scanning (D) is too late, and skipping scanning (E) is the risk to avoid.

Last updated Sep 18, 2026