Codex Path
Codex · Mock Exam 1
A 50-item, domain-weighted independent mock exam for the OpenAI Academy Codex pathway, with full explanations and a readiness readout. Not an official OpenAI assessment.
This is a full-length, domain-weighted independent mock exam for the Codex track. It is built from publicly available OpenAI learning objectives and is not an official OpenAI assessment, and not the Academy Codex assessments themselves. All 50 questions are new and do not repeat the domain-page items. Use this as your diagnostic: sit it first to find your two weakest domains.
Instructions
- Time: 60 minutes (our design choice for a 50-item mock; the Academy assessments are shorter, randomised selections from a larger bank).
- Items: 50, multiple-choice and multiple-response. Each item states how many answers to select.
- Selection: for multiple-response items you must select all correct options and no incorrect ones; partial selections are marked wrong.
- No guessing penalty: answer every question.
- Target: aim for at least 80% raw (≈ 40/50) before taking the real Academy Codex assessment, which passes at ≥ 80%.
- Work each question before expanding the answer.
Domain distribution
| # | Domain | Items here |
|---|---|---|
| 1 | Codex Fundamentals and Surfaces | 10 |
| 2 | Core Coding Workflows | 12 |
| 3 | Extending and Configuring Codex | 10 |
| 4 | Team Adoption and Governance | 10 |
| 5 | Scaling Across Teams and Systems | 8 |
Total: 10 + 12 + 10 + 10 + 8 = 50 items.
Readiness interpretation
This is an independent readiness indicator, not a score and not a prediction of any official result.
| Raw score (of 50) | Band | Interpretation |
|---|---|---|
| 45–50 | Strong readiness | 90%+; strong across all domains |
| 40–44 | Assessment ready | 80–89%; at or above the Academy threshold |
| 35–39 | Building confidence | 70–79%; close, target your weak domains |
| under 35 | Keep learning | below 70%; revisit the domain pages before re-attempting |
The 80% line is deliberate: it matches the Academy badge threshold.
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A developer keeps a chat-driven Codex session open on their laptop alongside other ChatGPT work. Which Codex surface are they using?
Show answer
Answer: B.
The ChatGPT desktop app embeds Codex in the desktop client for chat-driven coding alongside other ChatGPT work. The CLI (A) is a terminal agent, cloud (C) runs in a hosted sandbox away from the machine, and the IDE extension (D) lives inside an editor for a tight edit-review loop.
Which reasoning level in the Codex CLI automatically delegates a task to subagents running in parallel?
Show answer
Answer: C.
Ultra is the top rung and automatically delegates to parallel subagents. Max (A) gives more thinking time on a single task, and Extra High (B) and High (D) apply deeper reasoning on one task without parallel delegation.
Which model does OpenAI's guidance describe as the choice for the hardest end-to-end work needing sustained reasoning, judgment and multi-tool use?
Show answer
Answer: C.
GPT-6 Astra is guidance's pick for the hardest end-to-end work. Luna (A) is for high-volume mechanical tasks, Terra (B) is the pragmatic all-rounder, and Spark (D) is a text-only research preview for near-instant iteration.
A developer wants to launch a Codex session and set the model to GPT-5.6 at start-up from the shell. Which flag does that?
Show answer
Answer: B.
codex -m(equivalently--model) selects the model at launch.--effort(A) does not exist for model selection,/model(C) is an in-session command not a shell flag, and--reasoning(D) is not the model flag.The
gpt-5.3-codex-sparkmodel is best characterised as which of the following?Show answer
Answer: B.
Spark is a text-only research preview for near-instant iteration on ChatGPT Pro. It is not a default (A), the gpt-5.4 replacement is Terra/Luna (C), and it is text-only rather than a vision model (D).
After the 31 August 2026 changes, which pair of models replaces
gpt-5.4andgpt-5.4-miniin Codex under ChatGPT sign-in?Show answer
Answer: C.
Terra replaces
gpt-5.4and Luna replacesgpt-5.4-mini. GPT-5.5 (A) is a previous generation, Astra/Sol (B) are for harder work, andgpt-5.3-codex/gpt-5.2(D) are themselves deprecated under ChatGPT sign-in.Which statement about how Codex surfaces relate to configuration is accurate?
Show answer
Answer: B.
Codex is one agent behind many surfaces that share a single
config.toml. They are not separate agents (A), the CLI is not the only configurable surface (C), andAGENTS.mdholds per-repository guidance, not surface configuration (D).Which TWO situations point clearly toward using Codex cloud rather than the CLI or IDE extension?
Show answer
Answer: A and C.
Codex cloud runs in a hosted sandbox, ideal for long-running (A) and parallel (C) work independent of your machine. A quick edit in the editor (B) and a tight IDE loop (E) point to the IDE extension, and a scripted CI job (D) points to
codex execon the CLI.A service authenticates to Codex with an API key rather than ChatGPT sign-in. How does the 31 August 2026
gpt-5.4retirement affect it?Show answer
Answer: B.
The retirement applies to Codex with ChatGPT sign-in; API-key sign-in is explicitly unaffected. So the retirement is not universal (A), there is no earlier deadline for API keys (C), and switching to ChatGPT sign-in (D) would bring the retirement into scope, not avoid it.
In the CLI reasoning ladder, which order of increasing effort is correct?
Show answer
Answer: A.
The CLI ladder is Low, Medium, High, Extra High, Max, Ultra. Options B and C scramble the order, and 'Light' (D) is the GUI-client label, not the CLI ladder, and omits Extra High.
What is the single biggest lever on the quality of a Codex change before any reasoning effort is chosen?
Show answer
Answer: B.
A tightly scoped task drives output quality far more than model or effort. Model cost (A) will still sprawl on a vague task, the surface (C) does not fix scope, and prompt length (D) is not the same as scope.
Codex says 'I have fixed the bug and all tests pass.' What is the correct interpretation of that statement?
Show answer
Answer: B.
The model's claim is a lead, not evidence; you verify by running the tests and reading the diff. It is not proof (A), it says nothing about effort settings (C), and it is a useful lead worth checking rather than ignoring (D).
In the review-first loop, why ask Codex to produce a plan before it writes code?
Show answer
Answer: B.
The plan step is the cheapest place to catch a wrong approach, before any code exists. It is not a CLI requirement (A), does not choose the model (C), and is not primarily a token-saving measure (D).
A repository has no durable record of its build command, test command and 'do not touch' areas, so Codex keeps guessing. Where should these live?
Show answer
Answer: B.
AGENTS.md, checked into the repo, is the durable home for build/test commands and no-go areas. A source comment (A) is easy to miss, shell history (C) is per-machine and transient, and a PR description (D) is per-change, not durable context.Which command runs a scoped change once, non-interactively, on the low-cost model for a mechanical edit?
Show answer
Answer: B.
codex execis the non-interactive form and-m gpt-5.6-lunapicks the low-cost model for mechanical work.--interactive(A) is the opposite mode,codex plan(C) is not the execution form and Astra is overkill, andcodex chat(D) is not the non-interactive command.A developer wants a long refactor to run on its own branch, locally, without disturbing their current working tree. What fits best?
Show answer
Answer: A.
A git worktree checks a branch out into a separate directory so the task runs in isolation locally. Committing to main (B) is unsafe, deleting the tree (C) is destructive and unnecessary, and disabling tests (D) removes verification.
When reading a Codex-produced diff, which of the following most warrants closer scrutiny?
Show answer
Answer: C.
Unexpected edits to migrations or generated files are the risky, out-of-scope signals to scrutinise. Consistent formatting (A), reuse of utilities (B) and a good commit message (D) are neutral-to-positive.
A pull request says only 'Codex did this, tests pass.' Which TWO additions would make it reviewable in minutes?
Show answer
Answer: A and C.
Reviewable evidence is a focused diff (A) and a passing test run with a regression test (C). The model ID (B), the effort level (D) and a claim about the model (E) do not help a reviewer judge correctness or scope.
The goal is 'make the failing test payments.test.ts pass.' What is the cleanest definition of done to give Codex?
Show answer
Answer: B.
A failing test is an executable definition of done; requiring it to pass while the suite stays green scopes the task precisely. Prose only (A) is vaguer, a large speculative change (C) invites scope creep, and skipping the test (D) defeats the purpose.
A developer re-runs the same vague prompt at higher and higher reasoning effort with disappointing results. What is the most likely root cause?
Show answer
Answer: B.
Poor results from a vague prompt are usually a scoping problem, not an effort problem. A bigger model (A) still sprawls on a vague task, the CLI is not implicated (C), and repository size (D) is not the cause of vague output.
For a bug fix, why is reproducing the original reported case before and after important even when the suite is green?
Show answer
Answer: B.
A green suite that never ran the real repro can hide a lookalike fix; reproducing the reported case and adding a regression test confirms the specific fix. A green suite alone does not prove it (A), and reproduction relates to correctness, not model cost (C) or test speed (D).
Which TWO steps belong to the verification stage of a review-first workflow before merging a Codex change?
Show answer
Answer: A and B.
Verification means running the tests yourself (A) and reading the diff for scope creep and risky edits (B). Raising effort and re-running (C) does not verify anything, merging first (D) removes the gate, and deleting tests (E) destroys the evidence.
A team needs Codex to read from and write to an external ticketing system. Which extension point fits best?
Show answer
Answer: B.
MCP servers and connectors are how Codex reaches external systems and data. Hooks (A) run your own logic at lifecycle points, record and replay (C) captures sessions, and rules (D) constrain behaviour — none provide external-system access.
For an unattended
codex execjob in CI, which configuration is the safest?Show answer
Answer: B.
Unattended runs need a restrictive mode plus a sandbox so an agent cannot escalate silently. Broad auto-approve (A) is the risk to avoid, no permission model (C) is unsafe, and an interactive profile (D) is usually too permissive for CI.
How should auto-review be characterised in a safe workflow?
Show answer
Answer: B.
Auto-review reads the diff and summarises it as reviewer evidence, but does not replace the human gate. It does not replace reviewers (A), does not auto-merge (C), and is not limited to formatting (D).
A team on a ChatGPT Enterprise workspace wants to enable experimental context management so the model keeps notes across a long task. What is true at launch?
Show answer
Answer: B.
Experimental context management is opt-in and, at launch, limited to Plus/Pro sign-in — not Business, Enterprise or API-key. It is not universal (A), not on by default (C), and not tied to Luna (D).
On Windows, which TWO options let Codex run in a contained environment?
Show answer
Answer: A and C.
Windows sandbox and WSL are the supported contained environments for Codex on Windows. Force-pushing (B) and committing to a protected branch (E) are risky git actions, and disabling permissions (D) removes containment rather than adding it.
A team wants to run their own policy check automatically before Codex writes any files. Which extension point fits best?
Show answer
Answer: A.
Hooks run your own logic at lifecycle points such as before a write. A skill (B) is a packaged ability Codex invokes, record and replay (C) captures sessions, and a larger model (D) does not enforce a policy check.
What is the safe default for internet access on a Codex cloud task?
Show answer
Answer: B.
Cloud internet access should default to restricted and be opened deliberately for the minimum a task needs. Always-open (A) is over-permissive, access is controllable (C), and it is not gated to API-key sign-in (D).
Which distinction between skills and plugins is correct?
Show answer
Answer: B.
Skills are packaged, invokable abilities, while plugins bundle behaviour and are typically managed at team scale. They are not identical (A), external access is MCP not skills (C), and plugins are not cloud-only (D).
Which TWO items belong in a repository's AGENTS.md rather than the shared config.toml?
Show answer
Answer: A and C.
AGENTS.mdcarries per-repository build/test commands (A) and conventions and no-go areas (C). The default model (B) and a workspace-wide permission profile (D) areconfig.toml/workspace-level, and secrets (E) do not belong in a checked-in guidance file.Why should a restrictive permission mode and a sandbox be used together for a risky run rather than either alone?
Show answer
Answer: B.
Permission modes govern approvals while the sandbox contains impact; together they limit both what happens without a human and how far any action reaches. They are not redundant (A), and neither alone is sufficient (C, D).
An admin wants consistent model availability, permissions and allowed extensions across 100 engineers. What is the best mechanism?
Show answer
Answer: B.
Managed configuration lets an admin set and enforce settings centrally, avoiding drift. Self-configuration (A) and a document (C) rely on individuals matching settings by hand, and defaults (D) do not enforce the team's policy.
A CI pipeline running in GitHub Actions needs to authenticate to Codex without storing a long-lived secret. Which option fits best?
Show answer
Answer: B.
Workload identity federation lets cloud and CI workloads authenticate without a stored long-lived secret. A committed PAT (A) and an env-file password (D) are stored secrets, and a shared admin password (C) is both a secret and a governance failure.
Which identity is most appropriate for a shared, org-owned automation bot?
Show answer
Answer: B.
A service account is a non-human, org-owned identity, which is what a shared automation should use. A personal account (A) or the admin's PAT (D) ties the automation to an individual, and a group inbox (C) is not an auth identity.
A repository handles protected health information. What configuration is required?
Show answer
Answer: B.
Workloads handling protected health information require HIPAA configuration. 'Internal' (A) does not exempt PHI, a stricter model (C) is not the compliance control, and disabling Codex (D) is unnecessary when HIPAA configuration exists.
A compliance team needs an auditable record of Codex actions for investigations. Which surface provides it?
Show answer
Answer: B.
The Compliance API and audit events provide the action record for compliance and investigation. The analytics dashboard (A) and Analytics API (C) are for adoption and usage, and
AGENTS.md(D) is repo guidance.Which TWO are reliable ways to report Codex adoption to leadership?
Show answer
Answer: A and C.
Workspace analytics (dashboard) and the Analytics API (programmatic) are the real measurement surfaces. Anecdotes (B), shell history (D) and license-count guesses (E) are not reliable adoption measures.
Where should security scanning of Codex-produced code run so it happens before every merge?
Show answer
Answer: B.
Integrating Codex Security into CI or GitLab CI ensures code is scanned before merge automatically. A manual step (A) gets skipped, post-incident scanning (C) is too late, and trusting the model without scanning (D) is the risk to avoid.
A team gives every engineer admin rights to reduce permission errors. What is the governance problem?
Show answer
Answer: B.
Universal admin breaks least privilege and widens the blast radius of mistakes and compromise. It is not simply efficient (A), admins can code (C), and it does not disable analytics (D).
What controls which plugins, connectors and skills a workspace may use?
Show answer
Answer: B.
Workspace-level plugin/connector/skill controls decide which extensions are allowed.
AGENTS.md(A) is per-repo guidance, local settings (C) are per-engineer and ungoverned, and the model picker (D) selects models, not extensions.Which TWO authentication choices correctly match their actor in a Codex rollout?
Show answer
Answer: A and C.
Workload identity federation fits a CI/cloud workload without a stored secret (A), and a service account fits an org-owned automation bot (C). A PAT for a shared bot (B) ties it to a person, the admin's personal account for everything (D) destroys accountability, and a shared admin password (E) is a security and governance failure.
Two parallel Codex tasks overwrote each other's changes on one feature branch. What is the correct fix?
Show answer
Answer: B.
Isolating each workstream on its own branch or worktree and merging through a controlled gate prevents collisions while keeping parallelism. Running fewer tasks (A) sacrifices throughput, merging to main (C) is unsafe, and disabling tests (D) removes verification.
Codex behaves inconsistently across 20 repositories with different build commands and conventions. What is the best remedy?
Show answer
Answer: B.
Standardising
AGENTS.mdand configuration makes agent behaviour predictable across repos and makes parallel outputs safe to integrate. Accepting drift (A) leaves the problem, a bigger model (C) does not fix inconsistent context, and universal admin (D) is a governance error.You want to add Codex steps to a GitHub CI workflow. Which integration fits best?
Show answer
Answer: B.
The GitHub Action embeds Codex into a GitHub CI workflow. The SDK (A) is for building Codex into your own tooling, and Slack (C) and Linear (D) are team and issue integrations, not CI steps.
Leadership reports Codex adoption tripled and concludes scaling is a success. What is the flaw in that conclusion?
Show answer
Answer: B.
Rising activity without a quality view can hide falling outcomes; adoption must be paired with quality metrics. More usage is not automatically better (A), adoption is measurable (C), and model size (D) is unrelated to the measurement flaw.
Which TWO practices make parallel Codex workstreams safe to integrate?
Show answer
Answer: A and C.
Isolation per branch or worktree and independent verification before a controlled merge are the safe-integration practices. Unreviewed merges (B), a shared branch (D) and turning off CI (E) remove the safeguards that make parallelism safe.
A team wants to build Codex capability into their own internal developer tool. Which integration fits best?
Show answer
Answer: B.
The Codex SDK gives programmatic control to build Codex into your own tools. The GitHub Action (A) is for GitHub CI, Slack (C) is a team surface, and record and replay (D) captures sessions rather than embedding capability.
A team on GitLab plans to rely on the Codex GitLab integration for a critical launch next week. What should they weigh?
Show answer
Answer: B.
The GitLab integration is in beta, which matters when planning a critical launch. It is not GA-equivalent to GitHub (A), GitLab is supported in beta rather than unsupported (C), and migrating to GitHub (D) is not required.
What is the correct relationship between adoption metrics and quality metrics when scaling Codex?
Show answer
Answer: B.
Adoption and quality are complementary: usage without quality can mask declining outcomes, so both are tracked. Adoption alone does not prove value (A), quality does not replace adoption measurement (C), and both are measurable (D).
Last updated Sep 18, 2026