# Governance & Security Checklist (OpenAI)

A layered deployment checklist for OpenAI in the enterprise — identity and access, network, data residency and retention, safety, operations, Codex-specific controls, and a sign-off matrix, built from published OpenAI documentation.

import { Steps } from '@prosefly/astro-components';

This is a layered checklist for deploying OpenAI products safely in an organisation, drawn from the published OpenAI documentation on enterprise controls, safety and Codex administration. It is independent preparation, not an official compliance framework — verify every control against the current OpenAI docs and your own regulatory obligations before you rely on it.

:::caution[Prompts do not enforce policy]
The recurring assessment trap is a stem where the tempting answer is a stronger system-prompt rule. Prompts guide behaviour; **identity, network, data and safety controls enforce it.** When a stem describes a data-leak, access, or residency risk, the discriminator is almost always a platform control, not a prompt.
:::

## The layers, at a glance

```text
        ┌───────────────────────────────────────┐
        │ Identity & access  (who can call)      │
        ├───────────────────────────────────────┤
        │ Network            (from where)        │
        ├───────────────────────────────────────┤
        │ Data               (what is stored)    │
        ├───────────────────────────────────────┤
        │ Safety             (what is allowed)   │
        ├───────────────────────────────────────┤
        │ Operations         (limits & audit)    │
        ├───────────────────────────────────────┤
        │ Codex              (agent-specific)    │
        └───────────────────────────────────────┘
     Each layer fails independently; secure all of them.
```

## Layer 1 — Identity and access

Control who can authenticate and what each identity may do.

- [ ] **SSO** — SAML SSO for the workspace so access follows your identity provider, not standalone passwords.
- [ ] **SCIM** — automated provisioning and, critically, de-provisioning so a leaver loses access immediately.
- [ ] **RBAC** — roles and workspace permissions granted least-privilege; not everyone needs admin.
- [ ] **Service accounts** — non-human identities for automation, separate from individual users, so a person leaving does not break a pipeline.
- [ ] **Workload identity federation** — for automated callers, federate from your platform (Kubernetes, AWS, Azure, GCP, OCI, GitHub Actions, SPIFFE, X.509) instead of long-lived static keys.
- [ ] **Personal access tokens (PATs)** — scoped and rotated; treat a leaked PAT as an incident; prefer federation or service accounts for automation over personal tokens.
- [ ] **Domain verification** — verify your domains so workspace membership maps to your organisation.

| Identity type | Use for | Avoid |
| --- | --- | --- |
| SSO user | Interactive human access | Shared logins |
| Service account | Automation and pipelines | Tying automation to a person's account |
| Workload identity federation | CI, clusters, cloud jobs | Long-lived static API keys in those contexts |
| Personal access token | Narrow personal automation | Broad, un-rotated, shared tokens |

## Layer 2 — Network

Control where calls originate and how traffic is protected.

- [ ] **Private Link** — private connectivity so traffic does not traverse the public internet where offered.
- [ ] **IP allowlist** — restrict access to known corporate egress ranges.
- [ ] **Mutual TLS (mTLS)** — mutual authentication for high-assurance service-to-service calls.
- [ ] **IP egress ranges** — pin the platform's published egress ranges in your firewall rules for inbound webhooks and callbacks.
- [ ] **Terraform provider** — manage the above as code so network posture is reviewable and reproducible, not click-configured.

## Layer 3 — Data

Control what is stored, where, and for how long.

- [ ] **Residency region** — choose a data-residency region that meets your obligation. ChatGPT enterprise data residency is available in ten regions: US, EU, UK, JP, CA, KR, SG, IN, AU, UAE.
- [ ] **No-training default** — confirm that business and enterprise data is **not used to train models by default**; this is the default for those tiers.
- [ ] **Retention** — set retention to the minimum the use case needs; log and review.
- [ ] **Zero Data Retention (ZDR)** — where available and required, so eligible request content is not retained.
- [ ] **Encryption** — TLS 1.2 in transit and AES-256 at rest are the platform baseline; EKM (enterprise key management) where you must hold keys.
- [ ] **Your-data controls** — apply the documented data controls for what is stored, shared and exported.

:::caution[Where ZDR does not apply]
The **Agents API managed harness is US data residency only and does not offer ZDR**, and using a self-hosted sandbox does not change that. If a workload requires ZDR or non-US residency, the managed Agents harness is the wrong choice — that is a common discriminator. Choose the Agents SDK or Responses-API-plus-tools with your own execution instead.
:::

## Layer 4 — Safety

Control what the system is allowed to produce and do.

- [ ] **Safety classifiers** — apply the documented safety classifiers to inputs and outputs.
- [ ] **Moderation** — screen user content and model output for policy violations before it reaches users or systems.
- [ ] **Red teaming** — probe the deployment adversarially before and after launch; document findings and fixes.
- [ ] **Misalignment monitoring** — monitor for behaviour drifting from intended objectives, especially in agentic deployments.
- [ ] **Content provenance** — apply provenance signals for generated content where disclosure matters.
- [ ] **Under-18 guidance** — follow the documented under-18 guidance for any deployment reachable by minors.
- [ ] **CSAM guidance** — follow the documented CSAM guidance; this is non-negotiable and reportable.
- [ ] **Cybersecurity checks** — apply the documented cybersecurity checks; the security-specialised models are gated to authorised security work only.

## Layer 5 — Operations

Control cost, throughput and auditability.

- [ ] **Rate limits** — set per-project rate limits so one workload cannot starve another.
- [ ] **Spend limits** — set spend limits and alerts so a runaway job does not become a runaway bill.
- [ ] **Admin API** — manage the workspace, members, projects and keys programmatically.
- [ ] **Analytics API** — pull usage analytics for cost attribution and adoption reporting.
- [ ] **Compliance API and audit events** — export audit events to your SIEM so every administrative and access action is logged and reviewable.
- [ ] **Error handling** — handle documented error codes with backoff on transient errors; do not retry 4xx blindly.

## Layer 6 — Codex-specific controls

Codex acts on code, so it needs agent-specific governance beyond the layers above.

- [ ] **Permission modes and profiles** — set the permission mode per environment; the agent should not have blanket write or execute rights.
- [ ] **Sandboxing** — run agent execution sandboxed (including Windows sandbox and WSL where relevant); network access controlled.
- [ ] **Cloud internet-access controls** — restrict what a cloud Codex environment can reach.
- [ ] **Approvals** — require agent approvals and human review for sensitive actions; do not auto-approve destructive operations.
- [ ] **Auto-review and Codex Security** — enable auto-review; use Codex Security (plugin, CLI, cloud) for scans, deep scans, the security workbench, triage, fixes and vulnerability reports, wired into CI or GitLab CI.
- [ ] **Managed configuration** — enforce a managed `config.toml` and `AGENTS.md` rules so individuals cannot silently widen permissions.
- [ ] **Auth for agents** — use workload identity, service accounts or scoped personal access tokens; note that some ChatGPT-sign-in models retire on schedule (for example `gpt-5.4` and `gpt-5.4-mini` retire from Codex with ChatGPT sign-in on 31 August 2026) while API-key sign-in is unaffected.
- [ ] **Groups, provisioning and lifecycle** — manage Codex users through groups, provisioning and a user-lifecycle process, with roles and workspace permissions.
- [ ] **Sector controls** — apply HIPAA configuration and integrations such as Prisma AIRS where the sector requires it.

## Who signs off on what

A control is only real if someone owns it. This matrix is a starting point; adapt it to your org.

| Control area | Proposes | Reviews | Approves / owns |
| --- | --- | --- | --- |
| SSO / SCIM / RBAC | Platform engineering | Security | IT / Identity lead |
| Service accounts & federation | Platform engineering | Security | Security lead |
| Network (Private Link, allowlist, mTLS) | Network engineering | Security | Network / Security lead |
| Data residency & retention | Data owner | Legal / Privacy | DPO / Privacy officer |
| ZDR / EKM requirements | Product owner | Legal / Security | DPO + Security lead |
| Safety (classifiers, moderation, red team) | Applied ML / Trust & Safety | Security | Trust & Safety lead |
| Under-18 / CSAM guidance | Trust & Safety | Legal | Legal + Trust & Safety |
| Rate & spend limits | Product owner | Finance | Engineering manager |
| Audit export (Compliance API) | Platform engineering | Security | Security / Compliance lead |
| Codex permissions & sandboxing | Engineering lead | Security | Engineering manager + Security |

## A minimal go-live sequence

<Steps>
1. **Identity first** — SSO and SCIM live, RBAC least-privilege, automation on service accounts or federation, no shared logins.
2. **Data decisions recorded** — residency region chosen, no-training default confirmed, retention set, ZDR or EKM decided where required (and the Agents-harness residency constraint noted).
3. **Network fenced** — allowlist, Private Link or mTLS as the assurance level demands, managed as code.
4. **Safety wired in** — moderation and classifiers on inputs and outputs, red-team findings closed, under-18 and CSAM guidance applied where reachable.
5. **Operations instrumented** — rate and spend limits set, audit events flowing to the SIEM, error handling with backoff.
6. **Codex governed** — permission modes, sandboxing, approvals, auto-review and managed config in place before agents touch a real repo.
7. **Sign-off recorded** — every row in the matrix has a named owner and a dated approval.
</Steps>

## Common misconceptions

| Misconception | Reality | Why it matters on the assessment |
| --- | --- | --- |
| "A system prompt keeps data safe" | Data safety is residency, retention, ZDR and access, not prompts | Prompt-as-enforcement trap |
| "Business data is used to train by default" | Business and enterprise data is not trained on by default | Data-default trap |
| "The managed Agents harness can do ZDR / EU residency" | It is US-only, no ZDR; self-hosted sandbox does not change it | Residency trap |
| "Personal access tokens are fine for CI" | Prefer workload identity federation or service accounts | Identity trap |
| "Codex can auto-approve everything to move fast" | Sensitive and destructive actions need approvals and sandboxing | Agent-governance trap |
| "Moderation is optional if the prompt is careful" | Classifiers and moderation are enforcement layers | Safety-layer trap |

## Key takeaways

- Secure every layer independently: identity, network, data, safety, operations, and Codex — a gap in one undoes the others.
- Enforce with platform controls (access, residency, retention, moderation, sandboxing), never with prompt text.
- Prefer SSO, SCIM, RBAC, service accounts and workload identity federation over long-lived personal keys.
- Confirm no-training-by-default and choose residency, retention, ZDR and EKM to match your obligation.
- The Agents managed harness is US-only with no ZDR; route ZDR or non-US workloads elsewhere.
- Govern Codex as an agent: permission modes, sandboxing, approvals, auto-review, and a managed configuration.
- Every control needs a named owner and a dated sign-off, and audit events must reach your SIEM.
