AI Cert Prep
Type to search documentation.

Codex Path

D3 · Extending and Configuring Codex

The shared config.toml and AGENTS.md, rules, subagents, skills, plugins, MCP, hooks, record and replay, permission modes and profiles, sandboxing, auto-review, cloud internet access and experimental context management.

This domain is worth 20% of the mock — roughly 10 of 50 items. It tests whether you can configure Codex safely and extend it deliberately: the shared config.toml and AGENTS.md, the extension points (skills, plugins, MCP, hooks), and above all the safety surface — permission modes and profiles, sandboxing, auto-review and internet-access controls. Many items reward the safe configuration, not the most permissive one.

What you need to know

One shared config.toml configures the desktop app, CLI and IDE extension; AGENTS.md carries per-repository agent guidance, and rules and subagents shape behaviour. Codex is extended through skills, plugins, MCP servers, hooks, record & replay, and site tools (WebMCP). Its safety surface is where the exam concentrates: permission modes and profiles, sandboxing (including Windows sandbox and WSL), auto-review of diffs, and cloud internet access controls. There is also an opt-in experimental context management feature with specific sign-in limits.

Learning objectives

By the end of this page you should be able to:

  1. Configure Codex through the shared config.toml and per-repo AGENTS.md, rules and subagents.
  2. Extend Codex with skills, plugins, MCP servers and hooks, and know what each adds.
  3. Choose permission modes and profiles appropriate to the risk of a task.
  4. Apply sandboxing — including Windows sandbox and WSL — to contain what Codex can do.
  5. Use auto-review and cloud internet-access controls as part of a safe workflow.
  6. Explain experimental context management and its sign-in limits.

3.1 The shared config.toml

One configuration file serves the desktop app, the CLI and the IDE extension. Set the model, and reference the config basics/advanced/reference docs, environment variables and the sample config for the full surface.

toml
# config.toml — shared by desktop app, CLI and IDE extension
model = "gpt-5.6"
[features]
# opt-in features live here (see 3.7)
Configuration layerWhat it setsScope
config.tomlModel default, features, permission/profile settingsAll surfaces on this machine/workspace
Environment variablesOverrides and secretsProcess/session
AGENTS.mdPer-repository agent guidanceThe repository
RulesBehavioural constraints on the agentAs configured

Assessment signal

Set the model once for the CLI, IDE and desktop app → the shared config.toml. Per-repository build commands and conventions → AGENTS.md. Do not conflate the two: config is machine/workspace-wide defaults; AGENTS.md is repo-specific guidance.

3.2 AGENTS.md, rules and subagents

Agent configuration shapes how Codex works in a given repository.

  • AGENTS.md — the durable, checked-in guidance: build/test commands, conventions, no-go areas (covered in D2).
  • Rules — behavioural constraints that steer or forbid actions.
  • Subagents — delegation of parts of a task to parallel agents (the mechanism behind the Ultra reasoning level).
  • Speed — configuration that trades latency against thoroughness.
text
config.toml AGENTS.md rules subagents
(machine/ (per-repo (behavioural (parallel
workspace build, conventions, constraints) delegation)
defaults) no-go areas)
└───────── how Codex is set up ──────────┘ └── how work is split ──┘

3.3 Extending Codex: skills, plugins, MCP, hooks, record & replay

Codex has several distinct extension points; the exam expects you to tell them apart.

Extension pointWhat it addsReach for it when
SkillsReusable, packaged capabilities Codex can invokeYou want to give Codex a repeatable, named ability
PluginsBundled extensions to Codex’s behaviourYou are adding managed functionality across a team
MCPModel Context Protocol servers and connectorsYou need Codex to reach an external tool or data source
HooksRun your own logic at lifecycle pointsYou want to gate, log or transform steps of a run
Record & replayCapture and re-run a sessionYou want reproducibility or to share a run
Site tools (WebMCP)Web-based tool accessYou need web tooling exposed to the agent

Assessment signal

Reach an external system / data source → MCP. Run my own check at a lifecycle point (before write, after test) → hooks. A packaged, reusable ability → skills. Reproduce or share a session → record & replay.

3.4 Permission modes and profiles

Permission modes govern what Codex may do without asking. Profiles bundle settings for a context (personal, team, CI). This is the highest-value safety lever in the domain.

text
LESS PERMISSIVE ───────────────────────────────► MORE PERMISSIVE
ask before every auto-approve reads, broad auto-approve
sensitive action prompt on writes / (only in a trusted,
(read-only leaning) network / escalation sandboxed context)
Match the mode to the RISK of the task, not to convenience.
  • Use a restrictive mode for unattended runs (codex exec in CI) so an agent cannot escalate silently.
  • Use profiles to keep a safe default for shared machines and a looser one only inside a sandbox.
  • Approvals are the human decision point for escalations — treat them as a review gate, not a nuisance.

3.5 Sandboxing, including Windows sandbox and WSL

Sandboxing contains what Codex can touch — the filesystem, the network, the shell. Environments include local, cloud and git worktrees; on Windows, the Windows sandbox and WSL are supported.

EnvironmentWhat it containsNotes
Local sandboxFilesystem and command scope on your machinePair with a restrictive permission mode
Cloud sandboxA hosted environmentUsed by Codex cloud; internet access is controlled separately (3.6)
Git worktreeBranch isolationKeeps long-running work off your main tree
Windows sandboxIsolated Windows environmentSupported for Windows workflows
WSLWindows Subsystem for LinuxSupported for Linux-toolchain work on Windows

Assessment signal

Contain what the agent can touch, run it isolated, on Windows → sandboxing (Windows sandbox / WSL). Combine sandbox + restrictive permission mode for unattended or risky runs.

3.6 Auto-review and cloud internet access

Two features complete the safe-workflow picture.

  • Auto-review — Codex reviews the diff it produced and surfaces a summary you can attach as reviewer evidence (see D2 §2.6). It is an aid to review, not a replacement for a human reviewer.
  • Cloud internet access controls — govern whether a cloud run can reach the network. Default to no or restricted access; open it deliberately when a task genuinely needs it (e.g., fetching a dependency), and remember the data-handling implications.
text
Diff produced ─► AUTO-REVIEW summary ─► attach as evidence ─► HUMAN review ─► merge
(agent's own read; (D2 §2.6) (still required)
not a substitute)

3.7 Experimental context management

Astra can keep notes across context windows and search earlier messages and tool results within the same task. This is experimental and opt-in.

toml
[features.context_management]
experimental_mode = true

Sign-in limits matter: at launch it is available on ChatGPT Plus/Pro sign-in only — not Business, Enterprise or API-key sign-in. So an enterprise team on a Business/Enterprise workspace cannot rely on it yet, and an item that offers it as the fix for an enterprise team is wrong on availability alone.

Assessment signal

Keep notes across context windows, search earlier tool results in the same task → experimental context management, opt-in via features.context_management.experimental_mode. Enterprise / Business / API-key wants it → not available to them at launch.

Decision framework

Use CONFIGURE → CONTAIN → EXTEND → REVIEW (CCER) when setting Codex up for a task or team.

StepQuestionThe move
ConfigureWhat is the default model and per-repo guidance?Set model in config.toml; write AGENTS.md; add rules
ContainWhat must the agent be prevented from doing?Choose a permission mode/profile and a sandbox matched to the risk
ExtendWhat capability or data does the task need?Add skills/plugins for abilities, MCP for external systems, hooks for lifecycle logic
ReviewHow will the change be checked?Turn on auto-review for evidence; keep a human review gate; control cloud internet access

Common mistakes

MistakeWhy it happensWhat to do instead
Running unattended work with broad auto-approveFewer prompts feels fasterUse a restrictive permission mode + sandbox for codex exec / CI
Confusing config.toml with AGENTS.mdBoth configure CodexConfig is machine/workspace defaults; AGENTS.md is per-repo guidance
Treating auto-review as a full reviewIt reads the diff for youAuto-review is evidence; a human review gate is still required
Reaching for a plugin when you need external dataPlugins sound generalExternal systems/data are MCP; plugins bundle behaviour
Leaving cloud internet access wide openThe task once needed the netDefault to restricted; open access deliberately per task
Expecting experimental context management on EnterpriseIt is the newest featureIt is Plus/Pro sign-in only at launch; not Business/Enterprise/API-key
No sandbox on a risky runTrusting the modelContain the run; combine sandbox and permission mode
Using a hook where a rule fits (or vice versa)The extension points overlapHooks run your logic at lifecycle points; rules constrain agent behaviour

Scenario challenge

Scenario. A fintech team wants Codex to run a nightly job that upgrades dependencies and opens a PR. They also want the agent to “keep notes across the long task so it does not lose context”, and they run on a ChatGPT Enterprise workspace. An engineer proposes: broad auto-approve so the job never stalls, full cloud internet access, experimental context management turned on, and auto-review as the only review before the PR merges automatically.

Expert reasoning trace.

  1. Permission mode is wrong. Broad auto-approve on an unattended nightly job means the agent can escalate silently — exactly the risk to avoid. The safe configuration is a restrictive permission mode plus a sandbox, with any escalation surfaced rather than auto-approved.
  2. Internet access is over-open. A dependency upgrade may need network access to fetch packages, but full internet access is more than the task requires and carries data-handling risk in a fintech context. Grant the minimum the task needs, deliberately.
  3. Experimental context management is unavailable here. At launch it is Plus/Pro sign-in only; a ChatGPT Enterprise workspace cannot use it. The proposal fails on availability, so “keep notes across the long task” must be solved another way (scoping, subagents, or accepting the limit) — not by this feature.
  4. Auto-review is not a substitute for human review. It produces a useful summary to attach as evidence, but auto-merging on auto-review removes the human gate. For a fintech dependency change, a human review before merge is exactly the control you keep.
  5. Reassemble the safe design. Restrictive permission mode + sandbox, minimal deliberate internet access, drop the experimental feature on this workspace, and require a human review of the auto-review summary and diff before merge.

Exam-correct decision: restrictive permission mode with a sandbox, least-privilege cloud internet access granted only for the fetch, no reliance on experimental context management on an Enterprise workspace, and auto-review used as evidence feeding a required human review gate — not auto-merge. Not broad auto-approve, not full internet access, not experimental context management on Enterprise, not auto-merge on auto-review.

Assessment traps

TrapWhy it is temptingThe discriminator
“Broad auto-approve so the job never stalls”Fewer interruptionsUnattended runs need a restrictive mode + sandbox; escalations must surface
“Auto-review replaces the reviewer”It reads the diffAuto-review is evidence; a human gate is still required
“Turn on experimental context management for our Enterprise team”It is the newest capabilityPlus/Pro sign-in only at launch; not Business/Enterprise/API-key
“Use a plugin to reach the ticketing system”Plugins sound like integrationsExternal systems/data are MCP; plugins bundle behaviour
“Set repo build commands in config.toml”It is the config fileBuild commands and conventions belong in AGENTS.md
“Full internet access is simplest”It removes a variableLeast privilege: grant only what the task needs, deliberately
“A hook and a rule are the same”Both shape behaviourHooks run your logic at lifecycle points; rules constrain the agent

Practice questions

Each item states how many responses to select. Attempt before revealing.

Q1 · Which file configures the model default for the desktop app, CLI and IDE extension at once? (Select one)

A. A separate settings file per surface B. The shared config.toml C. AGENTS.md D. The PR template

Answer: B. Codex surfaces share one config.toml, where model = '…' sets the default for all of them. Per-surface files (A) contradict the shared design, AGENTS.md (C) holds per-repo guidance, and a PR template (D) is unrelated.

Q2 · A team needs Codex to read from and write to an external issue tracker. Which extension point fits BEST? (Select one)

A. A hook B. Record & replay C. An MCP server / connector D. A rule

Answer: C. MCP servers and connectors are how Codex reaches external systems and data sources. Hooks (A) run your logic at lifecycle points, record & replay (B) captures sessions, and rules (D) constrain behaviour — none of them provide external-system access.

Q3 · For an unattended `codex exec` job in CI, which permission configuration is safest? (Select one)

A. Broad auto-approve so it never stalls B. A restrictive permission mode plus a sandbox, so escalations surface and the blast radius is contained C. No permission model at all D. Whatever the developer uses interactively

Answer: B. Unattended runs need a restrictive mode and a sandbox so the agent cannot escalate silently and its actions are contained. Broad auto-approve (A) is the risk to avoid, no permission model (C) is unsafe, and an interactive profile (D) is usually too permissive for CI.

Q4 · What is the correct characterisation of auto-review? (Select one)

A. It replaces human review entirely B. It produces a review summary of the diff that you can attach as evidence, while a human review gate is still required C. It automatically merges the PR D. It only checks formatting

Answer: B. Auto-review reads the diff and produces a summary that serves as reviewer evidence, but it does not replace the human gate. It does not replace reviewers (A), does not auto-merge (C), and is not limited to formatting (D).

Q5 · An Enterprise-workspace team wants to enable experimental context management so Astra keeps notes across a long task. What is true? (Select one)

A. It is available to all sign-in types B. At launch it is ChatGPT Plus/Pro sign-in only and is not available on Business, Enterprise or API-key sign-in C. It is enabled by default D. It only works with gpt-5.6-luna

Answer: B. Experimental context management is opt-in and, at launch, limited to Plus/Pro sign-in — not Business, Enterprise or API-key. It is not universal (A), not on by default (C), and not tied to Luna (D).

Q6 · Which TWO belong in `AGENTS.md` rather than the shared `config.toml`? (Select two)

A. The repository’s build and test commands B. The machine-wide default model C. Per-repository ‘do not touch’ paths and coding conventions D. The workspace’s global permission profile E. Secrets shared across all repos

Answer: A and C. AGENTS.md carries per-repository build/test commands and conventions and no-go areas. The default model (B) and a workspace-wide permission profile (D) are config.toml/workspace-level, and secrets (E) do not belong in a checked-in guidance file.

Q7 · A team wants to run their own logic — a policy check — automatically before Codex writes any files. Which extension point fits BEST? (Select one)

A. A hook at the relevant lifecycle point B. A skill C. Record & replay D. A larger model

Answer: A. Hooks run your own logic at lifecycle points, such as before a write. A skill (B) is a packaged ability Codex invokes, record & replay (C) captures sessions, and a larger model (D) does not enforce a policy check.

Q8 · On Windows, which options let Codex run in a contained environment? (Select two)

A. Windows sandbox B. Force-pushing to main C. WSL (Windows Subsystem for Linux) D. Disabling all permissions E. Committing directly to a protected branch

Answer: A and C. Windows sandbox and WSL are the supported contained environments for Codex on Windows. Force-pushing (B) and committing to a protected branch (E) are risky git actions, not sandboxes, and disabling permissions (D) removes containment rather than adding it.

Q9 · For a Codex cloud task, what is the safe default for internet access? (Select one)

A. Always fully open, for convenience B. Restricted by default, opened deliberately only for the specific need such as fetching a dependency C. Internet access cannot be controlled D. Only available on API-key sign-in

Answer: B. Cloud internet access should default to restricted and be opened deliberately for the minimum the task needs. Always-open (A) is over-permissive, access is controllable (C), and it is not gated to API-key sign-in (D).

Q10 · What is the correct way to enable experimental context management? (Select one)

A. It is on automatically for Astra B. Opt in via features.context_management.experimental_mode = true in config.toml, subject to the Plus/Pro sign-in limit C. Pass --context on every command D. Enable it in AGENTS.md

Answer: B. It is opt-in through features.context_management.experimental_mode = true, and only on Plus/Pro sign-in at launch. It is not automatic (A), not a per-command flag (C), and not configured in AGENTS.md (D).

Q11 · A developer wants to reproduce and share an exact Codex session for debugging. Which feature fits BEST? (Select one)

A. Skills B. Record & replay C. MCP D. Auto-review

Answer: B. Record & replay captures a session so it can be re-run or shared. Skills (A) are reusable abilities, MCP (C) provides external access, and auto-review (D) summarises a diff — none reproduce a session.

Q12 · Why should a restrictive permission mode and a sandbox be used together for a risky run? (Select one)

A. They are redundant; either one alone is enough B. The permission mode decides what needs approval; the sandbox contains the blast radius if something does run — together they limit both the decision surface and the impact C. Only the sandbox matters D. Only the permission mode matters

Answer: B. Permission modes govern approvals while the sandbox contains impact; combining them limits both what happens without a human and how far any action can reach. They are not redundant (A), and neither alone is sufficient (C, D).

Q13 · A team confuses skills and plugins. What is the clearest distinction to give them? (Select one)

A. They are identical B. Skills are reusable, named capabilities Codex can invoke; plugins are bundled extensions to Codex’s behaviour, often managed across a team C. Skills reach external systems; plugins do not exist D. Plugins are only for the cloud surface

Answer: B. Skills are packaged, invokable abilities, while plugins bundle behaviour and are typically managed at team scale. They are not identical (A), external access is MCP not skills (C), and plugins are not cloud-only (D).

Q14 · An engineer sets repository build commands and coding conventions in `config.toml` and is surprised another repo does not use them. What is the correction? (Select one)

A. config.toml is per-repo; the second repo needs its own B. Per-repository build commands and conventions belong in that repo’s AGENTS.md; config.toml holds machine/workspace-wide defaults C. Conventions cannot be configured D. Both files must contain identical content

Answer: B. Repo-specific build commands and conventions go in AGENTS.md; config.toml is for machine/workspace defaults, which is why the other repo did not inherit them. config.toml is not per-repo (A), conventions are configurable (C), and the files serve different purposes so need not match (D).

Key takeaways

  • One shared config.toml sets defaults (model, features, permissions) for the desktop app, CLI and IDE extension; AGENTS.md holds per-repository guidance.
  • Extend deliberately: skills for reusable abilities, plugins for bundled behaviour, MCP for external systems/data, hooks for lifecycle logic, record & replay for reproducibility.
  • Match permission modes and profiles to the risk; unattended runs need a restrictive mode plus a sandbox.
  • Sandboxing contains the blast radius; on Windows, use the Windows sandbox or WSL.
  • Auto-review produces reviewer evidence but does not replace a human review gate.
  • Default cloud internet access to restricted; open it deliberately for the specific need.
  • Experimental context management is opt-in via features.context_management.experimental_mode and, at launch, is Plus/Pro sign-in only — not Business/Enterprise/API-key.

Last updated Sep 18, 2026