AI Cert Prep
Type to search documentation.

Domains

D5 · Governance, Safety & Risk Management

Layered guardrails, LLM failure-mode taxonomy, human-in-the-loop design, regulatory compliance (GDPR, HIPAA, FedRAMP, SOC 2, EU AI Act), data retention and ZDR, ethical AI, model risk management, incident response, and red-teaming.

This domain is worth 14% – roughly 9 of 63 items. It tests whether you can make a Claude system safe, compliant and defensible in a regulated enterprise. The recurring judgements: enforce safety in layers (not a single prompt rule), put human gates on irreversible/regulated actions, and match data handling to the applicable regulation (GDPR/HIPAA/FedRAMP/SOC 2/EU AI Act) rather than convenience.

Learning objectives

By the end of this page you should be able to:

  1. Design layered guardrails (defence in depth).
  2. Reason with an LLM failure-mode taxonomy.
  3. Place human-in-the-loop gates and design approval UX and audit.
  4. Map deployments to regulatory frameworks: GDPR (DSAR, DPIA, residency), HIPAA (BAA, PHI), FedRAMP High, SOC 2, EU AI Act.
  5. Choose data-retention posture, including Zero Data Retention and Fable 5.1’s 30-day requirement.
  6. Apply ethical-AI principles (bias, fairness, transparency, explainability, disclosure).
  7. Run risk assessment / model risk management, incident response, and red-teaming within the acceptable-use policy.

5.1 Layered guardrails (defence in depth)

No single control is sufficient. Safety comes from independent layers, so a bypass of one is caught by the next.

text
input ─▶ [1 input classifier] ─▶ [2 system-prompt rules] ─▶ [3 tool-permission hooks]
─▶ Claude ─▶ [4 output validation] ─▶ [5 human review] ─▶ action
(prompt-injection (soft guidance, (HARD enforcement of (schema/safety (irreversible/
& PII screen) not enforcement) business rules) checks) regulated gate)
LayerEnforcesStrength
Input classifierScreen prompt injection, PII, disallowed contentProbabilistic
System-prompt rulesBehavioural guidanceSoft — can be bypassed
Tool-permission hooksHard business/authorization rulesDeterministic (exit code 2 blocks)
Output validationSchema, safety, policy checksDeterministic
Human reviewJudgement on high-stakes/irreversibleStrongest, slowest

Prompt is not a guardrail

A system-prompt sentence is layer 2 — guidance, not enforcement. Hard rules (spend limits, deletions, data access) belong in hooks and validation (anti-pattern #3). Any answer that relies solely on prompt wording to enforce a critical rule is wrong.


5.2 LLM failure-mode taxonomy

Failure modeWhat it isPrimary control
HallucinationFluent, unsupported contentGrounding, citations, validation
Prompt injection (direct)User overrides instructionsInput classifier, privileged/instruction separation
Prompt injection (indirect)Malicious instructions in tool results/docs/webTreat external content as untrusted; sanitise; least privilege
JailbreakBypass safety policyLayered classifiers, refusal behaviour
Data exfiltrationCoaxing out secrets/PIIPII minimisation, output filtering, no secrets in context
Excessive agencyAgent does more than intendedLeast privilege, human gates, idempotency
SycophancyAgrees with the user regardless of truthNeutral prompting, independent review

Exam signal

“A document/tool result contains hidden instructions” → indirect prompt injection; the fix is to treat external content as untrusted data and apply least privilege, not to trust it because it came from a tool.


5.3 Human-in-the-loop design

Human review is mandatory for irreversible, regulated, external, or personal-data-bearing actions. The architect decides where the gate goes, what the reviewer sees, and how it is audited.

Design questionGuidance
Where does the gate go?Immediately before the irreversible/regulated action (payment, deletion, publish, clinical/legal output)
What does the reviewer see?The proposed action, the evidence/citations, and the confidence — enough to judge, not a rubber stamp
How is escalation triggered?On task complexity/risk, not on sentiment and not on self-reported confidence
How is it audited?Immutable log of who approved what, when, with the inputs and model version

Escalation anti-patterns

Do not escalate based on sentiment (sentiment ≠ complexity, anti-pattern #5) or self-reported confidence (anti-pattern #4). Escalate on measured complexity/risk signals or deterministic rules.


5.4 Regulatory compliance

FrameworkCore obligationsArchitecture implications
GDPRLawful basis; DSAR (access/erasure/rectification); DPIA for high-risk processing; data residencyEU-region placement (Vertex/Bedrock EU); deletion path; minimise/retention limits; DPIA before launch
HIPAABAA with the provider; protect PHI; minimum necessarySign a BAA; PHI handling controls; ZDR/retention posture; audit logs
FedRAMP HighUS federal authorised environmentDeploy via Bedrock/Vertex in FedRAMP High boundary (not direct API)
SOC 2Security/availability/confidentiality controls, auditedInherit provider controls; document your own
EU AI ActRisk-tiered obligations; transparency; high-risk system dutiesClassify the system’s risk tier; disclosure; human oversight; documentation

Exam signal

“US federal / FedRAMP High” → Bedrock or Vertex in the FedRAMP boundary, never direct API. “Healthcare / PHI” → BAA + PHI controls + retention. “EU personal data / right to erasure” → GDPR: DPIA, EU residency, deletion path.


5.5 Data retention and Zero Data Retention

PostureWhat it meansConstraint
Standard retentionProvider may retain inputs/outputs per policy for a periodDefault
Zero Data Retention (ZDR)Inputs/outputs not retained beyond serving the requestAvailable for eligible models/enterprise — not for Fable 5.1
Fable 5.1Requires 30-day data retention; not eligible for ZDR; not in Priority TierIf a workload mandates ZDR, do not use Fable 5.1

Fable 5.1 vs a ZDR requirement

If a compliance requirement demands ZDR (e.g. certain regulated data), Fable 5.1 is disqualified because it requires 30-day retention. Choose Opus 5 / Sonnet 5 with ZDR instead. This is a direct exam trap.


5.6 Ethical AI

PrinciplePractice
Bias & fairnessStratified evals per protected group; symmetric treatment; human review of impactful decisions
TransparencyDocument data sources, model versions, and known limitations
ExplainabilityCite sources; show reasoning where appropriate; make decisions auditable
DisclosureTell users they are interacting with AI where required; label AI-generated content
Human dignityKeep consequential decisions (hiring, credit, clinical) under human authority

5.7 Risk assessment and model risk management (MRM)

Treat the AI system as a managed risk, not a black box.

MRM elementPractice
Risk tieringClassify by blast radius, reversibility, regulation, data sensitivity
Controls mappingMap each risk to a guardrail layer / human gate
Model inventoryTrack model IDs, versions, prompts, eval scores, owners
Change controlRe-run regressions and re-assess risk on any model/prompt change
MonitoringOnline metrics, drift detection, incident triggers

5.8 Incident response for AI systems

Triggers: safety-metric spike, exfiltration alert, injection detection, downstream complaint, cost anomaly. Correlation IDs and traces (D3) make an incident reconstructable.


5.9 Acceptable use and red-teaming

  • Operate within Anthropic’s Usage Policy / acceptable-use terms; prohibited uses are out of scope regardless of technical feasibility.
  • Red-teaming: proactively attack your own system — prompt injection (direct and indirect), jailbreaks, exfiltration, excessive-agency probes — before adversaries do. Feed findings into the regression suite and guardrail layers.
  • Secrets live in env / secret managers, never in prompts, CLAUDE.md, or logs. Log without secrets or PII.

5.10 Compliance control mapping (technical controls per framework)

The exam tests whether you can map a regulation to the specific architecture control it demands, not just name it.

Requirement in the stemFrameworkConcrete architecture control
“EU residents’ personal data”, “right to erasure”GDPREU-region placement (Vertex/Bedrock/Foundry EU); DSAR/erasure pipeline; DPIA before launch; retention limits; lawful basis
“high-risk processing of personal data”GDPRMandatory DPIA as a design-gate exit criterion
“patient records”, “PHI”, “clinical”HIPAASigned BAA; minimum-necessary access; audit logs; encryption; retention/ZDR posture
“US federal agency”, “FedRAMP High”FedRAMP HighDeploy via Bedrock/Vertex inside the authorised boundary; not direct API
“SOC 2 report”, “auditor evidence”SOC 2Inherit provider controls; document your own access/change/monitoring controls
“AI system classified high-risk”, “transparency to users”EU AI ActRisk-tier classification; user disclosure; human oversight; technical documentation and logging
“no data retained beyond serving the request”ZDRZDR-eligible model (Opus 5 / Sonnet 5) — never Fable 5.1 (30-day retention)

Where each control lives in the design

text
Design gate ──▶ DPIA (GDPR) · BAA (HIPAA) · risk-tier (EU AI Act) · residency + placement decision
Build gate ──▶ retrieval-time ACL/tenant filter · PII scrub · audit logging · encryption · secret manager
Runtime ──▶ ZDR/retention posture · human gate on irreversible actions · output validation
Ops ──▶ breach-notification runbook (GDPR 72 h) · incident response · immutable approval log

Exam signal

Match the signal word to the control: “FedRAMP” → Bedrock/Vertex boundary; “PHI” → BAA + PHI controls; “erasure/EU residents” → GDPR DPIA + residency + deletion path; “ZDR mandated” → not Fable 5.1. Answers that satisfy the regulation with a prompt rule or convenience choice are wrong.


5.11 A risk register for a Claude deployment

A risk register is the artefact that turns “it might go wrong” into owned, mitigated risk. Score likelihood × impact, assign an owner, and map each risk to a guardrail layer.

IDRiskLikelihoodImpactMitigation (layer)OwnerResidual
R1Indirect prompt injection via fetched docsMedHighUntrusted-content handling; least privilege; output validationSec engLow
R2Cross-tenant data leak in shared indexLowHighRetrieval-time tenant_id/ACL filter; cache keyed by tenantPlatformLow
R3Over-privileged agent runs destructive toolMedHighRemove tools (least privilege); permission hook; human gateArchLow
R4PHI mishandledLowHighBAA; minimum-necessary; audit logs; ZDR postureComplianceLow
R5Model deprecation breaks a pinned serviceMedMedModel inventory; regression suite; canary migrationArchLow
R6Aggregate metric masks a segment failureHighMedPer-segment evals + regression gateMLLow
R7Runaway agent loop / cost spikeMedMedstop_reason control + iteration backstop; cost alertsSRELow

Residual risk is communicated, not hidden

The register’s job is to make residual risk explicit to compliance and sponsors (via the risk register artefact, D6). A design that pretends risk is zero after mitigation is not credible.


5.12 Scenario walkthrough: a regulated healthcare intake assistant

Scenario. A hospital group wants a Claude assistant that reads patient intake forms (PHI), suggests a triage category, and drafts a note for the clinician. EU-based (GDPR + HIPAA-equivalent obligations apply via their US arm), on AWS. A vendor proposes Fable 5.1 “for the best reasoning”, auto-triage without clinician sign-off “to save time”, a single system-prompt rule “never expose PHI”, and indefinite retention “for quality”.

Expert reasoning trace.

  1. Retention/model first. PHI plus a regulated posture typically mandates ZDR or strict retention — which disqualifies Fable 5.1 (30-day retention). Choose Sonnet 5 / Opus 5 with a ZDR-eligible posture.

  2. Placement. EU residency + existing AWS → Bedrock in an EU region; sign a BAA for PHI.

  3. Human gate. Triage that affects care is consequential → the clinician retains authority; auto-triage without sign-off is rejected.

  4. PHI is not protected by a prompt. “Never expose PHI” as prose is prompt-as-enforcement. Enforce with minimum-necessary access, output validation, audit logging, and retrieval-time ACLs.

  5. Retention. Indefinite retention violates GDPR minimisation; set retention limits and a DSAR/erasure path. Run a DPIA at the design gate.

  6. Layered guardrails + register. Input PII screen → prompt guidance → tool hooks → output validation → clinician review; record R1–R7-style risks with owners and residual risk.

Why the tempting alternatives are wrong: Fable 5.1 fails the ZDR/retention requirement; auto-triage removes the mandatory human gate; a prompt rule can’t guarantee PHI protection; indefinite retention breaches GDPR minimisation and the erasure right.


5.13 Common misconceptions

MisconceptionRealityWhy it matters on the exam
“A strong prompt enforces a safety/business rule.”Prompts are layer-2 guidance; hard rules need hooks/validation.Prompt-as-enforcement is the top D5 wrong answer.
“ZDR is available on any model.”Fable 5.1 needs 30-day retention and is not ZDR-eligible.ZDR-mandate stems disqualify Fable 5.1.
“The direct API can serve FedRAMP High.”FedRAMP High is via Bedrock/Vertex boundaries.Federal stems require the cloud boundary.
“Internal tools don’t need compliance.”PHI/PII obligations apply regardless of internal use.“It’s internal” is never an exemption.
“Escalate when the user sounds upset.”Sentiment ≠ complexity/risk; escalate on measured signals.Sentiment-escalation is anti-pattern #5.
“Tool output is trustworthy because it’s a tool.”Fetched content can carry indirect injection; treat as untrusted.Indirect-injection stems test this.
“One strong guardrail is enough.”Defence in depth stacks independent layers.Single-layer answers are wrong.
“Full automation removes human bias.”It launders bias and removes accountability; keep human authority on consequential decisions.Hiring/credit/clinical stems require a human.

Exam traps in this domain

TrapWhy it is wrong
Enforcing a hard rule via the system prompt onlyPrompt is guidance (layer 2), not enforcement; use hooks/validation
Escalating to a human based on sentimentSentiment ≠ complexity (anti-pattern #5)
Escalating based on self-reported confidenceSelf-report is unreliable (anti-pattern #4)
Using direct API for a FedRAMP High workloadUse Bedrock/Vertex in the FedRAMP boundary
Using Fable 5.1 where ZDR is requiredFable 5.1 needs 30-day retention; disqualified
Trusting instructions embedded in a document/tool resultIndirect prompt injection; treat external content as untrusted
Processing EU personal data with no DPIA / EU residencyGDPR breach; DPIA + residency + deletion path required
Handling PHI without a BAAHIPAA violation; sign a BAA and apply PHI controls
Putting secrets in the prompt or CLAUDE.mdExfiltration risk; use secret managers
Single-layer guardrailNo defence in depth; one bypass = full failure
Skipping human review for irreversible/regulated actionsMandatory gate; confidence doesn’t remove it
No incident runbook / rollback pathCan’t contain or recover safely
Satisfying a regulation with a prompt rule or a convenience choiceRegulations demand specific technical controls, not prose
Treating residual risk as zero after mitigationResidual risk must be communicated in the risk register
Auto-executing consequential people-decisions (hiring/credit/clinical)Requires human authority and fairness evaluation
Missing the GDPR breach-notification timeline in the runbook72-hour notification is a legal obligation
Applying ACLs only in the prompt rather than at retrievalModel can be talked past prose; enforce at the data layer

Practice questions

Q1 · A financial agent must never move more than $10,000 without approval. The team plans to add 'never exceed $10,000 without approval' to the system prompt. What is the correct design? (Select one)

A. The prompt sentence is sufficient. B. Enforce the limit with a tool-permission hook / validation that blocks any transfer over $10,000 and routes it to a human approval gate; the prompt rule alone is not enforcement. C. Ask the model to report its confidence before transfers. D. Escalate only when the user sounds anxious.

Answer: B. Hard financial limits require deterministic enforcement (hook + human gate). The prompt is guidance only (A, anti-pattern #3). Self-reported confidence (C, #4) and sentiment-based escalation (D, #5) are all anti-patterns.

Q2 · A US federal agency requires FedRAMP High. Which deployment is compliant? (Select one)

A. Direct Anthropic API for newest models. B. Claude via Amazon Bedrock or Google Vertex AI within the FedRAMP High boundary. C. Any cloud, since Claude is inherently compliant. D. On a personal laptop.

Answer: B. FedRAMP High authorisation is provided through Bedrock/Vertex boundaries, not the direct API. The direct API (A) is not in the FedRAMP boundary; compliance is not inherent (C); a laptop (D) is absurd for federal data.

Q3 · A regulated workload mandates Zero Data Retention. The team wants Fable 5.1 for its reasoning. What is the correct guidance? (Select one)

A. Use Fable 5.1; ZDR applies to all models. B. Fable 5.1 requires 30-day data retention and is not ZDR-eligible, so choose a ZDR-eligible model (e.g. Opus 5 or Sonnet 5) for this workload. C. Turn off logging to achieve ZDR on Fable 5.1. D. Use Fable 5.1 but delete logs after 30 days.

Answer: B. Fable 5.1 mandates 30-day retention and cannot satisfy a ZDR requirement; a ZDR-eligible model is required. ZDR is not universal (A); disabling your own logging (C) doesn’t change provider retention; 30-day deletion (D) is exactly what ZDR forbids.

Q4 · An agent summarises web pages, and a page contains hidden text instructing it to email internal data to an external address. What is this, and the correct defence? (Select two)

A. Indirect prompt injection. B. Treat external content as untrusted; apply least privilege so the agent lacks an unrestricted email/exfiltration tool, and validate/deny such actions. C. Trust it because it came through a legitimate tool. D. Increase the model’s effort level. E. Add the instruction to the system prompt.

Answer: A and B. Malicious instructions in fetched content are indirect prompt injection; defence is untrusted-content handling plus least privilege and output/action validation. Trusting tool output (C) is the vulnerability; effort (D) is irrelevant; adding it to the prompt (E) executes the attack.

Q5 · A hospital wants a Claude assistant that reads patient records. What are the FIRST compliance requirements? (Select two)

A. A signed BAA with the model provider. B. PHI-handling controls (minimum necessary, access controls, audit logs) and an appropriate retention/ZDR posture. C. Nothing, since it is internal. D. Publishing patient data to improve the model. E. Escalating on patient sentiment.

Answer: A and B. PHI under HIPAA requires a BAA and PHI-handling controls with auditing and retention discipline. ‘Internal’ (C) does not exempt PHI; publishing patient data (D) is a violation; sentiment escalation (E) is an anti-pattern unrelated to compliance.

Q6 · A support system escalates to a human whenever the customer 'sounds angry'. Why is this flawed? (Select one)

A. It is optimal. B. Sentiment is not a proxy for complexity or risk; escalation should key off measured complexity/risk signals or deterministic rules, not emotion. C. It escalates too rarely. D. Anger always means the model failed.

Answer: B. Sentiment-based escalation (anti-pattern #5) conflates emotion with task difficulty; a calm customer can have a complex, high-risk case and vice versa. It is not optimal (A); the frequency (C) isn’t the core flaw; anger doesn’t imply model failure (D).

Q7 · Which set best represents layered guardrails (defence in depth)? (Select one)

A. A single, very detailed system prompt. B. Input classifier → system-prompt guidance → tool-permission hooks → output validation → human review for high-stakes actions. C. Only human review at the end. D. Only an output regex.

Answer: B. Defence in depth stacks independent layers so a bypass of one is caught by the next. A single prompt (A), review-only (C) or regex-only (D) are single points of failure.

Q8 · Processing EU residents' personal data with a high-risk profile, what must happen before launch? (Select two)

A. A DPIA (data protection impact assessment). B. EU data residency (e.g. Vertex/Bedrock EU region) and a data-subject deletion/erasure path. C. Nothing until a complaint arrives. D. Store all data indefinitely for quality. E. Escalate based on model confidence.

Answer: A and B. High-risk processing of EU personal data requires a DPIA and residency plus DSAR/erasure support under GDPR. Waiting for complaints (C) and indefinite storage (D) violate GDPR; confidence-based escalation (E) is an unrelated anti-pattern.

Q9 · A prompt-injection incident is detected in production. What is a sound first containment step? (Select one)

A. Delete all logs. B. Roll back to the previous prompt/model version and/or disable the offending tool via its permission hook, using correlation IDs and traces to scope the impact. C. Ignore it until the next release. D. Publicly disclose customer data to be transparent.

Answer: B. Containment means fast rollback and disabling the exploited capability, scoped via traces/correlation IDs. Deleting logs (A) destroys evidence; ignoring it (C) prolongs harm; disclosing customer data (D) causes a second breach.

Q10 · Why red-team a Claude system before launch? (Select one)

A. To satisfy marketing. B. To proactively find prompt-injection, jailbreak, exfiltration and excessive-agency weaknesses and feed fixes into guardrails and the regression suite before adversaries exploit them. C. Because red-teaming replaces evals. D. Because it guarantees zero risk.

Answer: B. Red-teaming surfaces vulnerabilities so they can be closed and turned into regression tests. It is not marketing (A), does not replace functional evals (C), and cannot guarantee zero risk (D).

Q11 · Where should secrets (API keys, DB credentials) live for a Claude agent? (Select one)

A. In the system prompt so the model can use them. B. In environment variables or a secret manager, never in prompts, CLAUDE.md, or logs. C. In the CLAUDE.md file for convenience. D. Hard-coded in the tool source and printed to logs.

Answer: B. Secrets belong in env/secret managers and must never enter prompts, config files the model reads, or logs. Putting them in the prompt (A), CLAUDE.md (C), or logs (D) creates exfiltration paths.

Q12 · A hiring-screening assistant makes consequential recommendations. Which governance controls are REQUIRED? (Select two)

A. Per-group fairness evaluation and symmetric treatment, reviewed for bias. B. A human decision-maker retains authority over the hiring decision (human-in-the-loop for a consequential, regulated decision). C. Fully automate decisions to remove human bias. D. Escalate only when a candidate sounds confident. E. Store all applicant data indefinitely.

Answer: A and B. Consequential decisions about people require fairness evaluation and human authority over the outcome. Full automation (C) launders bias and removes accountability; sentiment escalation (D) and indefinite retention (E) are anti-patterns/GDPR issues.

Q13 · A healthcare intake assistant handles PHI, is EU-based on AWS, and a vendor proposes Fable 5.1 with indefinite retention. Which TWO corrections are REQUIRED first? (Select two)

A. Choose a ZDR-eligible model (Opus 5 / Sonnet 5) because Fable 5.1’s 30-day retention fails a ZDR/PHI posture. B. Deploy on Bedrock in an EU region and sign a BAA, with retention limits and a DSAR/erasure path. C. Keep Fable 5.1 but promise to delete logs after 30 days. D. Rely on a system-prompt rule ‘never expose PHI’. E. Retain all data indefinitely for quality.

Answer: A and B. A PHI/ZDR posture disqualifies Fable 5.1 and requires EU residency, a BAA, retention limits and erasure. 30-day deletion (C) is exactly what ZDR forbids; a prompt rule (D) can’t protect PHI; indefinite retention (E) breaches GDPR minimisation.

Q14 · An architect must communicate residual risk (after mitigation) to a compliance stakeholder. Which artefact and content are correct? (Select one)

A. An ADR listing the technical decision only. B. A risk register scoring each risk by likelihood × impact with owner, mitigation layer, and the remaining residual risk. C. A cost model showing unit economics. D. A C4 code diagram.

Answer: B. A risk register with likelihood/impact, owners, mitigations and residual risk is the compliance-facing artefact. An ADR (A) records a technical decision; a cost model (C) is economics; a code diagram (D) is for engineers.

Q15 · A stem states 'high-risk processing of EU residents' personal data'. Which control belongs at the DESIGN gate? (Select one)

A. Wait for a complaint before assessing. B. Complete a DPIA and decide EU residency/placement so residency, retention and erasure shape the architecture before build. C. Add a disclaimer to the UI at launch. D. Store everything to be safe.

Answer: B. GDPR high-risk processing requires a DPIA and residency/erasure decisions at the design gate. Waiting (A) is too late; a UI disclaimer (C) doesn’t satisfy the DPIA; over-retention (D) breaches minimisation.

Q16 · A US federal agency mandates FedRAMP High and wants the newest model the day it ships. What is the correct guidance? (Select one)

A. Use the direct Anthropic API for earliest access. B. Deploy Claude via Bedrock or Vertex inside the FedRAMP High boundary; the residency/authorisation requirement overrides the desire for day-one access. C. Any cloud works; Claude is inherently authorised. D. Run it on a laptop in the office.

Answer: B. FedRAMP High is provided through the Bedrock/Vertex boundary, and the compliance requirement is binding over novelty. Direct API (A) is outside the boundary; authorisation isn’t inherent (C); a laptop (D) is unacceptable for federal data.

Q17 · An incident-response runbook for an EU personal-data system is missing one legally required element. Which is it? (Select one)

A. A marketing statement. B. The GDPR breach-notification step and 72-hour timeline to the supervisory authority. C. A list of favourite dashboards only. D. A cost projection.

Answer: B. GDPR requires breach notification (generally within 72 hours), so the runbook must include it. Marketing (A), a dashboard list (C), and a cost projection (D) are not the legal requirement.

Q18 · Which pair best maps signal words to the correct architecture control? (Select two)

A. ‘FedRAMP High’ → deploy via Bedrock/Vertex inside the authorised boundary. B. ‘ZDR mandated’ → do not use Fable 5.1 (30-day retention); use a ZDR-eligible model. C. ‘PHI’ → a strongly worded system prompt is sufficient. D. ‘EU erasure’ → retain data indefinitely. E. ‘high-risk AI’ → skip documentation.

Answer: A and B. FedRAMP maps to the cloud boundary and a ZDR mandate disqualifies Fable 5.1. PHI needs a BAA and controls, not a prompt (C); EU erasure requires a deletion path, not indefinite retention (D); high-risk AI (EU AI Act) requires documentation, not skipping it (E).

Key takeaways

  • Guardrails are layered: input classifier → prompt guidance → tool-permission hooks → output validation → human review. The prompt is guidance, never hard enforcement.
  • Know the failure taxonomy; treat external content (tool results, documents, web) as untrusted to counter indirect prompt injection.
  • Put human gates on irreversible/regulated actions; escalate on complexity/risk, never sentiment or self-reported confidence.
  • Match handling to the regulation: GDPR (DPIA, residency, erasure), HIPAA (BAA, PHI), FedRAMP High (Bedrock/Vertex), SOC 2, EU AI Act risk tiers.
  • Fable 5.1 requires 30-day retention and is not ZDR-eligible — disqualified where ZDR is mandated.
  • Practise ethical AI: fairness evals, transparency, disclosure, human authority over consequential decisions.
  • Run model risk management, an incident-response runbook with rollback, and red-teaming feeding the regression suite; keep secrets out of prompts, CLAUDE.md and logs.
  • Map signal words to controls: FedRAMP → Bedrock/Vertex boundary; PHI → BAA + controls; EU erasure/high-risk → DPIA + residency + deletion; ZDR → not Fable 5.1.
  • Maintain a risk register (likelihood × impact, owner, mitigation layer, residual) and communicate residual risk — never claim zero risk.
  • Put compliance controls at the right lifecycle stage: DPIA/BAA/risk-tier at design, ACL/PII/audit at build, retention/human-gate at runtime, breach-notification (72 h) in ops.

Last updated Sep 18, 2026