AI Cert Prep
Type to search documentation.

API Developer Path

D7 · Production Safety and Operations

Safety classifiers and moderation, red teaming, misalignment monitoring, content provenance, rate and spend limits, error codes and retries, RBAC, Private Link, IP allowlist, mTLS and workload identity federation.

This domain is about 8% of the OAI-API mock – roughly 5 of 60 items – and draws on the operations and safety material threaded through the Academy API pathway. It tests whether you can take a working application to production safely: moderating content, red-teaming and monitoring for misalignment, marking provenance, setting rate and spend limits, handling errors with retries, and locking down access with RBAC and network controls.

What you need to know

Production is where an application meets untrusted inputs, real money and real users. You screen content with moderation and safety classifiers, probe weaknesses with red teaming, and watch for misalignment in live behaviour. You mark AI-generated media with content provenance. You cap blast radius with rate limits and spend limits, and you handle transient failures with retries and exponential backoff keyed off the right error codes. You restrict who and what can reach the API with RBAC, Private Link, IP allowlists, mutual TLS and workload identity federation – so credentials are scoped, short-lived and network-bounded. A deployment checklist ties these together as launch gates.

Learning objectives

By the end of this page you should be able to:

  1. Apply moderation and safety classifiers to inputs and outputs.
  2. Explain red teaming and misalignment monitoring.
  3. Set rate limits and spend limits to bound blast radius.
  4. Handle error codes with correct retry and backoff behaviour.
  5. Choose access controls: RBAC, Private Link, IP allowlist, mTLS, workload identity federation.
  6. Run a deployment checklist before launch.

7.1 Safety: moderation, classifiers, red teaming, misalignment

ControlWhat it doesWhen
ModerationFlags disallowed content in inputs/outputsScreen user input and model output at runtime
Safety classifiersDetect specific risk categories (incl. cyber-safety checks)Gate high-risk flows
Red teamingDeliberate adversarial testing before and after launchFind jailbreaks, injection, data-exfiltration paths
Misalignment monitoringWatch live behaviour for drift from intended goalsOngoing, especially for agents with tools

Red teaming is proactive (you attack your own system); misalignment monitoring is continuous (you watch what it actually does). Content provenance (marking AI-generated media) supports transparency obligations.

Assessment signal

“Untrusted user input”, “jailbreak”, “prompt injection”, “before launch we should test adversarially” point to red teaming and guardrails/moderation. “Watch the deployed agent for drift” points to misalignment monitoring.

7.2 Rate limits and spend limits

Two different blast-radius controls:

LimitBoundsFailure it prevents
Rate limitRequests/tokens per minuteA bug or spike overwhelming the service; noisy-neighbour
Spend limitDollars per periodA runaway loop or abuse draining the budget
text
Runaway agent loops calling the API 1000×/min
│
rate limit → 429 caps the request rate
spend limit → hard cap on dollars, alerting before the ceiling

Set both before launch. A spend limit is the difference between a bug costing $50 and costing $50,000.

7.3 Error codes and retries

CodeMeaningCorrect handling
429Rate limit / quota exceededRetry with exponential backoff (respect Retry-After)
500 / 503Server / service errorRetry with backoff; it is transient
400Bad request (malformed)Do not retry; fix the request
401 / 403Auth / permissionDo not retry; fix credentials/scope
python
import time
for attempt in range(5):
try:
return client.responses.create(model="gpt-5.6-terra", input=q)
except RateLimitError: # 429
time.sleep(2 ** attempt) # 1, 2, 4, 8, 16s
except BadRequestError: # 400 — do not retry
raise

The discriminator on many items: retry transient errors (429/5xx) with backoff; never blindly retry a 400/401/403 – retrying a malformed or unauthorised request just wastes quota and can worsen rate limiting.

7.4 Access control and network security

ControlWhat it doesUse when
RBAC / permissionsScope who can do what (keys, roles, projects)Always; least privilege on keys and roles
Private LinkPrivate network path to the API, off the public internetSensitive workloads needing network isolation
IP allowlistOnly listed source IPs may call the APIFixed egress infrastructure
mutual TLS (mTLS)Both client and server authenticate with certsStrong mutual authentication requirements
Workload identity federationShort-lived federated credentials (K8s, AWS, Azure, GCP, OCI, GitHub Actions, SPIFFE, X.509) instead of long-lived keysCloud workloads that should not hold static secrets

Prefer federated, short-lived credentials

Long-lived API keys in environment variables are a standing liability. Workload identity federation lets a cloud workload exchange its own identity for a short-lived token, so there is no static key to leak. Prefer it wherever your platform supports it.

7.5 The deployment checklist

text
Before launch, confirm:
[ ] Moderation on untrusted inputs and outputs
[ ] Red-team pass for injection / jailbreak / exfiltration
[ ] Guardrails + human approval on irreversible actions
[ ] Rate limit set; spend limit set with alerting below the ceiling
[ ] Retry with exponential backoff on 429/5xx; no retry on 4xx auth/validation
[ ] RBAC least privilege on keys and roles
[ ] Network controls as required (Private Link / IP allowlist / mTLS)
[ ] Federated, short-lived credentials instead of static keys where possible
[ ] Data residency / retention verified against requirements (see D4)
[ ] Eval gate green (see D3); tracing/observability on (see D4)
[ ] Misalignment monitoring for deployed agents

Decision framework

Use GUARD to decide the controls a production launch needs.

LetterStepQuestion
GGate contentIs moderation on inputs and outputs, and have you red-teamed?
UUsage capsAre rate and spend limits set with alerting?
AAuth & accessIs RBAC least-privilege, with the right network controls?
RResilienceDo you retry transient errors with backoff and not retry auth/validation errors?
DDetect driftIs misalignment monitoring and tracing in place for agents?

Common mistakes

MistakeWhy it happensWhat to do instead
No spend limit“It won’t loop”Set a spend limit with alerting; it caps runaway cost
Retrying a 400/401Generic retry-everything wrapperRetry only transient 429/5xx with backoff; fix 4xx
Long-lived keys in env varsEasiest to set upUse workload identity federation for short-lived tokens
Moderating input but not outputFocus on user textScreen model outputs too, especially agent actions
Skipping red teaming“We tested the happy path”Adversarially test injection/jailbreak before launch
Over-broad API key scopeConvenienceRBAC least privilege per key/role/project
No misalignment monitoring for agentsLaunch-and-forgetContinuously watch deployed agent behaviour
Trusting the network by defaultPublic internet is fineAdd Private Link / IP allowlist / mTLS as risk requires

Scenario challenge

Scenario. You are launching an autonomous support agent that reads customer messages (untrusted input), calls internal tools to issue refunds, and runs on Kubernetes. Security asks: how do we stop it draining budget, stop a malicious message hijacking it, and avoid static API keys in the cluster? The team’s current plan wraps every API call in a “retry up to 5 times on any error” loop and stores a long-lived API key in a Kubernetes secret.

Expert reasoning trace.

  1. Budget blast radius → spend limit + rate limit. An agent that loops or is abused could call the API thousands of times. Set a spend limit with alerting below the ceiling and a rate limit, so a runaway loop hits a hard dollar cap and a 429 rather than an unbounded bill.
  2. Malicious message → red team + moderation + guardrails + approval. Customer messages are untrusted, so a prompt-injection payload could try to make the agent issue an unauthorised refund. Red-team for injection before launch, moderate inputs, add output/tool guardrails, and gate the refund action behind human approval (D4) since it is irreversible.
  3. Retry loop is wrong. “Retry any error 5×” will happily retry a 400 (malformed) and a 401/403 (auth/permission), wasting quota and masking real bugs. Correct: retry only 429/5xx with exponential backoff, and fail fast on 4xx auth/validation.
  4. Static key is a liability. A long-lived key in a K8s secret can leak. Use workload identity federation so the pod exchanges its Kubernetes identity for a short-lived token – no static key to steal.
  5. Network controls. If the workload’s egress is fixed, add an IP allowlist, and use Private Link if the traffic must stay off the public internet; RBAC scopes the credential to only the refund/support tools.
  6. Monitor after launch. Turn on misalignment monitoring and tracing so a drifting or hijacked agent is caught in production.

Exam-correct decision: spend + rate limits, red teaming and moderation with an approval gate on refunds, selective retry (429/5xx with backoff only), workload identity federation instead of a static key, RBAC and network controls, and misalignment monitoring. Not retry-everything, not a long-lived key in a secret, not an auto-refund with no gate.

Assessment traps

TrapWhy it is temptingThe discriminator
“Retry every error a few times”One wrapper for all failuresRetry only transient 429/5xx with backoff; never auth/validation 4xx
“A rate limit is enough to control cost”It limits requestsA rate limit bounds throughput; a spend limit bounds dollars – set both
“Store the API key in an env var / secret”SimpleUse workload identity federation for short-lived credentials
“Moderate the user input”User text feels like the riskScreen outputs and tool actions too, not just inputs
“We tested it works, we’re ready”Happy path passedRed-team injection/jailbreak/exfiltration before launch
“The agent is smart, no approval needed for refunds”Autonomy biasIrreversible actions need a human approval gate
“Public internet is fine for sensitive traffic”Default setupAdd Private Link / IP allowlist / mTLS to the risk level

Practice questions

Each item states how many responses to select. Commit before revealing.

Q1 · An API call returns a `429`. What is the correct handling? (Select one)

A. Retry immediately in a tight loop. B. Retry with exponential backoff, respecting any Retry-After header. C. Do not retry; fix the request. D. Escalate to a larger model.

Answer: B. A 429 is a transient rate-limit error, handled by backing off exponentially and honouring Retry-After. A tight retry loop (A) worsens the limiting, a 429 is not a malformed request to fix (C), and model size (D) is unrelated.

Q2 · Which error should you NOT automatically retry? (Select one)

A. 429 rate limit. B. 503 service unavailable. C. 400 bad request (malformed input). D. 500 server error.

Answer: C. A 400 means the request itself is malformed; retrying it will fail again and waste quota – fix the request instead. 429 (A), 503 (B) and 500 (D) are transient and appropriate to retry with backoff.

Q3 · A team fears a runaway agent loop could generate a huge bill. Which control MOST directly caps the dollar cost? (Select one)

A. A rate limit only. B. A spend limit with alerting below the ceiling. C. A larger context window. D. Streaming responses.

Answer: B. A spend limit caps dollars directly and can alert before the ceiling. A rate limit (A) bounds throughput, not total spend, and context size (C) and streaming (D) do not control cost.

Q4 · A Kubernetes workload should call the API without holding a long-lived key. What is the BEST approach? (Select one)

A. Store a long-lived key in a Kubernetes secret. B. Use workload identity federation to exchange the pod’s identity for a short-lived token. C. Hard-code the key in the container image. D. Pass the key as a command-line argument.

Answer: B. Workload identity federation issues short-lived, federated credentials so there is no static key to leak. A secret (A), a baked-in key (C) and a CLI argument (D) are all long-lived-key anti-patterns.

Q5 · An agent processes untrusted customer messages and can take actions. Which TWO controls address prompt-injection risk BEST? (Select two)

A. Input/output guardrails and moderation to detect and block injected instructions. B. Red teaming before launch to find injection and exfiltration paths. C. A larger context window. D. Higher reasoning effort. E. Removing all logging.

Answer: A and B. Guardrails/moderation block injected instructions at runtime and red teaming finds the paths before launch. Context size (C) and effort (D) do not stop injection, and removing logging (E) reduces the visibility you need.

Q6 · What is the difference between red teaming and misalignment monitoring? (Select one)

A. They are the same activity. B. Red teaming is proactive adversarial testing; misalignment monitoring is continuous observation of deployed behaviour for drift. C. Red teaming only applies to images. D. Misalignment monitoring happens only before launch.

Answer: B. Red teaming attacks the system to find weaknesses; misalignment monitoring watches live behaviour over time. They are distinct (A), red teaming is not image-only (C), and monitoring is ongoing after launch, not pre-launch only (D).

Q7 · A sensitive workload must reach the API without traversing the public internet. Which control fits? (Select one)

A. A spend limit. B. Private Link, providing a private network path to the API. C. A larger model. D. Predicted outputs.

Answer: B. Private Link gives a private network path off the public internet for sensitive traffic. A spend limit (A) is a cost control, and model size (C) and predicted outputs (D) are unrelated to network isolation.

Q8 · An API key used by a read-only reporting job also has permission to delete resources. What principle is violated and what is the fix? (Select one)

A. None; broad keys are convenient. B. Least privilege; scope the key via RBAC to only the read permissions the job needs. C. Statelessness; store nothing. D. Caching; enable it.

Answer: B. An over-scoped key violates least privilege and enlarges blast radius; RBAC should scope it to read-only. Convenience (A) is not a justification, and statelessness (C) and caching (D) are unrelated.

Q9 · Which items belong on a pre-launch deployment checklist for an agent that takes irreversible actions? (Select two)

A. A human approval gate on the irreversible actions. B. Rate and spend limits with alerting. C. Retrying every error type indefinitely. D. Disabling tracing to reduce noise. E. Storing static keys in the repo.

Answer: A and B. Irreversible actions need an approval gate, and rate/spend limits cap blast radius – both are launch gates. Retrying everything indefinitely (C) is wrong, disabling tracing (D) removes needed observability, and static keys in the repo (E) is a serious anti-pattern.

Q10 · A deployed support agent gradually starts taking actions outside its intended scope. Which capability is designed to catch this? (Select one)

A. Prompt caching. B. Misalignment monitoring of live agent behaviour, with tracing to investigate. C. Batch processing. D. A larger context window.

Answer: B. Misalignment monitoring watches deployed behaviour for drift from intended goals, and tracing lets you investigate. Caching (A) and Batch (C) are performance features, and context size (D) does not detect behavioural drift.

Key takeaways

  • Screen untrusted inputs and model outputs with moderation and safety classifiers; red-team before launch.
  • Set both a rate limit (throughput) and a spend limit (dollars) with alerting to bound blast radius.
  • Retry only transient errors (429, 5xx) with exponential backoff; never blindly retry 400/401/403.
  • Scope credentials with RBAC least privilege and add Private Link, IP allowlist or mTLS to the risk level.
  • Prefer workload identity federation and short-lived tokens over long-lived static keys.
  • Gate irreversible agent actions behind human approval and monitor deployed agents for misalignment.
  • Treat the deployment checklist as launch gates: safety, limits, resilience, access, residency, evals and observability.

Last updated Sep 18, 2026