AI Cert Prep
Type to search documentation.

CCAR-P Practice Exam 2

A second full-length, blueprint-weighted CCAR-P practice exam with new, harder, multi-constraint items.

This is a second full-length, 63-item practice exam for CCAR-P. Every item is new — none appears in Practice Exam 1 or on the domain pages — and the set is tuned to be slightly harder than exam 1: more multi-constraint stems, more capacity/cost arithmetic, and more FIRST / MOST cost-effective / TWO qualifiers that reward reading the whole stem before answering. The blueprint weighting is identical to exam 1, so your per-domain scores are directly comparable.

What is different from exam 1

  • Harder qualifiers. Several stems ask for the FIRST action or the MOST cost-effective design, so more than one option is defensible and you must pick the best.
  • Arithmetic items. You will compute per-request cost, RPM/ITPM/OTPM, cache savings, recall@k, MRR, and A/B significance from the numbers given.
  • Multi-constraint scenarios. Residency + existing cloud + irreversible action, or PHI + ZDR + placement, must all be satisfied at once.
  • Same distractor families as exam 1 (constraint-blind, over-engineered, prompt-as-enforcement, self-report, silent failure, recall-only, aggregate-metric, sentiment-as-complexity) so practising exam 1 transfers directly.

Domain distribution

#DomainWeightItems here
1Solution Design & Architecture17%11
2Claude Models, Prompting & Context Engineering13%8
3Integration (incl. RAG)19%12
4Evaluation, Testing & Optimization16%10
5Governance, Safety & Risk Management14%9
6Stakeholder Communication & Lifecycle Management14%9
7Developer Productivity & Operational Enablement7%4
Total100%63

How to use both exams

  1. Sit Exam 1 first, under timed conditions, before you touch exam 2. Use it to find your weak domains and restudy those domain pages.

  2. Use Exam 2 as the booking gate. Only after Exam 1 is comfortable (≥ 80% raw, no single domain far below the rest) should you sit this harder set. Treat a strong Exam 2 as the signal that you are ready to book.

  3. Compare per-domain results across both sittings. A domain that drops sharply from exam 1 to exam 2 is where the deeper, multi-constraint reasoning is still shaky — return to that domain page’s scenario walkthrough and misconceptions table.

Scoring proxy

The real exam scales 100–1000 with a pass at 720; there is no guessing penalty, so answer everything. As a raw proxy, aim for ≥ 80% (≈ 50/63) on this harder set before booking, and confirm no single domain is far below the others.

Take the practice exam

Interactive mode

Take the practice exam

63 questions · one at a time · 120-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · Solution Design and ArchitectureSelect one

    A workload runs 10 req/s average at 5k input + 1k output tokens on Sonnet 5 ($2/$10 per MTok) and must fit a $300k/month budget. Which design is the MOST cost-effective while preserving quality?

    • A. Run everything on Opus 5 for headroom.
    • B. Compute per-request cost (~$0.02), then layer prompt caching on the stable prefix, a cheap-first cascade escalating on validation failure, and Batch for tolerant traffic until the arithmetic clears the budget.
    • C. Guess a design and adjust after launch.
    • D. Move all traffic to Fable 5.1 for quality.
    Show answer

    Answer: B.

    Per-request cost is 5k×$2/1e6 + 1k×$10/1e6 = $0.02; cost-effectiveness is an arithmetic exercise of cache + cascade + batch layered until cost < budget. Blanket Opus (A) is over-engineered/overspends; guessing (C) is unanchored (constraint-blind); Fable 5.1 (D) is the most expensive model, the opposite of cost-effective.

  2. Q2D1 · Solution Design and ArchitectureSelect one

    Peak load is 25 req/s at 8k input tokens each; the tier caps ITPM at 6,000,000. What is the binding constraint and the FIRST sound response?

    • A. RPM is binding; nothing else matters.
    • B. ITPM is binding (peak ITPM = 25×60×8,000 = 12M > 6M); shift latency-tolerant traffic to Batch and/or raise the tier or spill overflow to a second model.
    • C. Send every request to Opus 5.
    • D. Ignore it because bursts are rare.
    Show answer

    Answer: B.

    peak_ITPM = 25×60×8,000 = 12,000,000, double the 6M cap, so ITPM binds first; the fix is Batch, a higher tier, or spillover. RPM-only (A) misreads the maths; Opus for all (C) raises cost and ITPM; ignoring bursts (D) invites a 429 storm (silent-failure/constraint-blind).

  3. Q3D1 · Solution Design and ArchitectureSelect one

    A p95 SLA of 5 s is being exceeded; traces show a long generated output dominates the latency budget. Which change is BEST?

    • A. Add a coordinator and three subagents for robustness.
    • B. Use a smaller/faster model or fast mode, shorten or stream the output, and lower effort where adequate, because generation is the dominant term.
    • C. Increase retrieval k to improve quality.
    • D. Move the interactive path to the Batch API.
    Show answer

    Answer: B.

    When generation dominates the latency budget, shrink model/output/effort or stream. Multi-agent (A) is over-engineered and adds latency; larger k (C) adds retrieval time; Batch (D) has up-to-24 h latency and cannot serve an interactive p95 SLA.

  4. Q4D1 · Solution Design and ArchitectureSelect two

    An EU-regulated insurer on Azure wants claims triaged with a fraud flag that must NEVER auto-deny a claim. Which TWO design choices are correct?

    • A. Deploy on Microsoft Foundry in an EU region for residency and existing IAM.
    • B. Route fraud auto-deny straight to the model to save adjuster time.
    • C. Insert a mandatory human approval gate before any denial, since it is irreversible and regulated.
    • D. Use the direct Anthropic API for the newest model.
    • E. Set a single aggregate accuracy target across all claim types.
    Show answer

    Answer: A and C.

    EU residency plus an existing Azure footprint point to Foundry EU, and an irreversible regulated action requires a human gate. Auto-deny (B) removes the mandatory gate; direct API (D) is constraint-blind to residency; a single aggregate target (E) hides per-claim-type failure (aggregate-metric).

  5. Q5D1 · Solution Design and ArchitectureSelect one

    During a 529 spike an active Fable 5.1 thinking session automatically falls back to a much older model and multi-turn behaviour becomes corrupted. What is the correct fix?

    • A. Raise the iteration cap.
    • B. Fall back only to a model with feature parity (equal-or-newer thinking support), log the degraded mode, and never silently downgrade a thinking session.
    • C. Disable retries so it fails fast.
    • D. Parse the output text to detect corruption.
    Show answer

    Answer: B.

    Thinking blocks are readable only by the producing model or newer, so fallback must preserve feature parity and log degraded mode. An iteration cap (A) and text parsing (D) don't address parity (prose-parsing anti-pattern); disabling retries (C) harms availability.

  6. Q6D1 · Solution Design and ArchitectureSelect one

    A team must launch in three weeks, lacks ops capacity, and the agent runtime is a commodity capability. What is the BEST build-vs-buy call?

    • A. Build a custom Agent SDK runtime and sandbox from scratch.
    • B. Use managed agents (Anthropic hosts the loop and sandbox) to hit the deadline with minimal ops.
    • C. Delay until an ML platform team is hired.
    • D. Buy a product that has no Claude support.
    Show answer

    Answer: B.

    No ops capacity plus a tight deadline plus a commodity capability points to managed/buy. Building (A) contradicts the constraints (over-engineered); delaying (C) fails the deadline (constraint-blind); a non-Claude product (D) abandons the requirement.

  7. Q7D1 · Solution Design and ArchitectureSelect one

    An agentic loop occasionally never terminates. Which is the CORRECT primary control plus backstop?

    • A. Scan the model's text for the word 'done' as the main signal.
    • B. Terminate on stop_reason (end_turn / no further tool_use), keeping an iteration cap only as a safety backstop.
    • C. Use a fixed 8-iteration cap as the sole control.
    • D. Lower temperature until it stops.
    Show answer

    Answer: B.

    Control flow keys off stop_reason; a cap is a backstop only. Text scanning (A) is the prose-parsing anti-pattern; a cap alone (C) is the iteration-cap-as-primary-stop anti-pattern; temperature (D) does not govern termination.

  8. Q8D1 · Solution Design and ArchitectureSelect one

    A stem describes independent parallel investigations that must each explore with isolated context and then be synthesised. Which pattern is justified?

    • A. A single augmented-LLM call.
    • B. Multi-agent orchestration (coordinator + subagents) with context isolation and fan-in synthesis.
    • C. A three-step sequential workflow sharing one context.
    • D. Fine-tuning on the research corpus.
    Show answer

    Answer: B.

    Parallel, context-isolated subtasks with fan-in are the textbook multi-agent signal. A single call (A) can't parallelise isolated exploration; a sequential shared-context workflow (C) doesn't isolate; fine-tuning (D) is unrelated to orchestration.

  9. Q9D1 · Solution Design and ArchitectureSelect one

    A design routes 100% to Opus 5 at 4× budget with acceptable quality. Which change is MOST cost-effective while preserving quality on hard cases?

    • A. Switch everything to Haiku 4.5 and accept lower accuracy.
    • B. Introduce a cheap-first cascade that escalates to Opus 5 only on a validation-check failure, plus prompt caching on the stable prefix.
    • C. Escalate on the model's self-reported confidence.
    • D. Reduce the number of users.
    Show answer

    Answer: B.

    A cascade routes cheap-first and escalates on validation failure, preserving quality on hard cases while most traffic runs cheaply; caching amortises the prefix. Blanket Haiku (A) sacrifices accuracy (constraint-blind); self-report (C) is anti-pattern #4; cutting users (D) doesn't address unit economics.

  10. Q10D1 · Solution Design and ArchitectureSelect one

    A stakeholder offers a vague goal and no numbers, and two teams disagree on scope. What is the BEST FIRST action?

    • A. Start building the most likely interpretation.
    • B. Run discovery to define measurable, per-segment success criteria and enumerate constraints, and get sign-off before designing.
    • C. Default to Opus 5 and begin.
    • D. Assume a target and proceed.
    Show answer

    Answer: B.

    Without criteria or constraints the design is unanchored; discovery surfaces the binding constraint and aligns stakeholders. Building (A), defaulting to a model (C), and inventing a target (D) all skip the anchoring step (constraint-blind).

  11. Q11D1 · Solution Design and ArchitectureSelect one

    Which change best completes an otherwise sound reference architecture that currently has no path from production signals back into the system?

    • A. Add a second load balancer.
    • B. Add a feedback loop: offline evals plus online metrics, traces and cost telemetry that inform prompt/model/retrieval changes.
    • C. Add more input validation only.
    • D. Increase the context window.
    Show answer

    Answer: B.

    A Professional-level architecture must close the loop from production signals into improvement; that is the feedback stage. A load balancer (A) is infrastructure; input validation (C) is the input stage; a bigger window (D) is unrelated to the missing loop.

  12. Q12D2 · Claude Models, Prompting and Context EngineeringSelect one

    A 4,000-token stable prefix on Sonnet 5 ($2/MTok input) is reused at a 90% cache-hit rate over 10,000 req/day. Approximately how much of input cost does caching save?

    • A. Nothing; caching never helps large prefixes.
    • B. About 70%, because a cache read is ~0.1× input so the 4k prefix behaves like ~400 tokens on the 90% of hits.
    • C. Exactly 50%, the Batch discount.
    • D. 100%; cached requests are free.
    Show answer

    Answer: B.

    Read ≈ 0.1× input turns 4,000 prefix tokens into ~400 on hits, dropping daily input cost from ~$90 to ~$27 (about 70%). Caching does help large stable prefixes (A); 50% is the Batch discount not caching (C); reads are cheap, not free (D).

  13. Q13D2 · Claude Models, Prompting and Context EngineeringSelect one

    A pipeline must return strictly typed JSON for downstream systems and runs on Fable 5.1, where forcing tool use returns 400. What is the BEST approach?

    • A. Force a tool call with tool_choice: "any" and retry on 400.
    • B. Use structured outputs with a strict JSON schema (or auto + instruction), guaranteeing the shape without forcing tool use.
    • C. Ask for JSON in the prompt and regex-parse the prose.
    • D. Switch to free-text output and reparse it.
    Show answer

    Answer: B.

    Structured outputs with a strict schema conform by construction and don't require forced tool use, which Fable 5.1 rejects with 400. Forcing tools then retrying (A) can't fix a 400; prompt-only JSON with regex (C) and free-text reparse (D) are brittle prose-parsing.

  14. Q14D2 · Claude Models, Prompting and Context EngineeringSelect one

    Caching is enabled but the hit rate is ~3%. Investigation shows a per-request request_id is prepended to the system prompt. What is the FIRST fix?

    • A. Disable caching; it doesn't work here.
    • B. Remove the volatile request_id from the prefix (log it separately) so the prefix is byte-stable, restoring cache hits.
    • C. Shorten the reference documents.
    • D. Increase the context window.
    Show answer

    Answer: B.

    A per-request token in the prefix changes it every call, so nothing caches; moving it out restores a stable prefix. Caching isn't broken (A); document length (C) only affects the minimum size; a bigger window (D) is unrelated.

  15. Q15D2 · Claude Models, Prompting and Context EngineeringSelect one

    A long agentic session on a thinking model overflows the window because tool results are huge. Which technique is correct, and what must be avoided?

    • A. Delete the earliest user and assistant turns client-side.
    • B. Use context editing to clear stale tool results server-side (and compaction for narrative), avoiding client-side edits to earlier reasoning that invalidate thinking blocks.
    • C. Lower the temperature to shrink the context.
    • D. Ask the model to summarise itself in the same turn.
    Show answer

    Answer: B.

    Context editing removes stale tool results server-side and compaction summarises narrative, both avoiding thinking-block invalidation. Deleting turns client-side (A) breaks the append-only invariant; temperature (C) doesn't affect context size; in-turn self-summary (D) doesn't reclaim the window.

  16. Q16D2 · Claude Models, Prompting and Context EngineeringSelect two

    A team runs every request on Opus 5 at effort: xhigh with acceptable quality and 5× budget; cost, not latency, is the complaint. Which TWO changes best cut cost while preserving quality?

    • A. Cascade most traffic to Sonnet 5, escalating to Opus 5 only on a validation-check failure.
    • B. Lower effort to an adequate level and re-run the per-segment regression suite to confirm no regression.
    • C. Escalate on the model's self-reported confidence.
    • D. Move all traffic to Fable 5.1 for maximum quality.
    • E. Remove the eval suite to cut compute.
    Show answer

    Answer: A and B.

    A cheap-first cascade and right-sized effort (re-validated per segment) cut cost while preserving quality on hard cases. Self-report (C) is anti-pattern #4; Fable 5.1 (D) is the most expensive model; removing evals (E) removes the quality guard (silent failure).

  17. Q17D2 · Claude Models, Prompting and Context EngineeringSelect one

    A hard rule 'never grant a discount above 20%' must be guaranteed in a Claude sales agent. Where does it belong?

    • A. A strongly worded sentence in the system prompt.
    • B. A programmatic tool-permission hook / validation that rejects any discount over 20% before execution.
    • C. A few-shot example refusing a large discount.
    • D. The thinking budget.
    Show answer

    Answer: B.

    Hard business rules require deterministic enforcement the model cannot be talked past. Prompt wording (A) and few-shot (C) are prompt-as-enforcement (anti-pattern #3); the thinking budget (D) is unrelated to enforcement.

  18. Q18D2 · Claude Models, Prompting and Context EngineeringSelect one

    After upgrading from Opus 4.6 to Opus 5, requests setting budget_tokens for thinking now return 400. What is the fix?

    • A. Add more retries.
    • B. Replace budget_tokens with thinking: {"type":"adaptive"} and control depth via effort, since budget_tokens is removed on Opus 5.
    • C. Downgrade permanently to Haiku 4.5.
    • D. Remove thinking entirely.
    Show answer

    Answer: B.

    budget_tokens is removed on Opus 5 (and Sonnet 5 / Fable 5.x) and returns 400; adaptive thinking plus effort is the supported replacement. Retries (A) can't fix a 400; Haiku (C) is the only model still using budget_tokens but is not an Opus equivalent; removing thinking (D) discards needed reasoning.

  19. Q19D2 · Claude Models, Prompting and Context EngineeringSelect one

    A durable fact (a customer's contract tier) must persist across separate sessions without re-sending it in every prompt. Which mechanism fits, and what constraint applies?

    • A. Paste the fact into every system prompt.
    • B. Use the memory tool to persist the fact across sessions, but never store secrets/PII there inappropriately and keep it out of model-visible logs.
    • C. Store it in CLAUDE.local.md.
    • D. Raise provider retention to keep it in logs.
    Show answer

    Answer: B.

    The memory tool persists durable facts across sessions; the constraint is not to store secrets/PII inappropriately. Pasting per prompt (A) bloats context; CLAUDE.local.md (C) is a Claude Code dev file, not a runtime store; relying on provider retention (D) is not a memory mechanism and raises compliance risk.

  20. Q20D3 · Integration (incl. RAG)Select one

    A B2B assistant serves 300 tenants from one shared vector index; a tenant occasionally sees another tenant's document in an answer. What is the correct control?

    • A. Add a system-prompt rule telling the model not to reveal other tenants' data.
    • B. Apply a mandatory tenant_id (and ACL) filter at retrieval time, before/alongside the vector search, so only authorised chunks ever enter context.
    • C. Filter the answer after generation to remove other tenants' data.
    • D. Give each tenant a larger model.
    Show answer

    Answer: B.

    Isolation is a retrieval-time security boundary: filter by tenant_id/ACL before the vector search. A prompt rule (A) is bypassable prompt-as-enforcement; post-generation filtering (C) is too late because the chunk already entered context; a bigger model (D) doesn't isolate data.

  21. Q21D3 · Integration (incl. RAG)Select one

    Across 5 queries the relevant chunk ranked 1, 4, 2, not-retrieved, and 1. What are recall@3 and MRR?

    • A. recall@3 = 1.00; MRR = 1.00.
    • B. recall@3 = 0.60; MRR = 0.55.
    • C. recall@3 = 0.55; MRR = 0.60.
    • D. recall@3 = 0.80; MRR = 0.70.
    Show answer

    Answer: B.

    Three of five relevant chunks are in the top 3 (ranks 1, 2, 1) so recall@3 = 3/5 = 0.60; reciprocal ranks 1, 0.25, 0.5, 0, 1 give MRR = 2.75/5 = 0.55. Option A ignores the misses; C swaps the values; D is arithmetically wrong.

  22. Q22D3 · Integration (incl. RAG)Select one

    recall@10 = 0.63 (low) and answers frequently omit the needed fact. Where should the architect work FIRST?

    • A. Grounding; tighten the answer-only-from-context instruction.
    • B. Retrieval: fix chunking/embeddings/hybrid and add reranking, because low recall means the relevant chunk often isn't retrieved at all.
    • C. Add more citations.
    • D. Switch the generation model to Opus 5.
    Show answer

    Answer: B.

    Low recall means the right chunk isn't reaching the top-k, so the failure is upstream in retrieval. Grounding and citations (A, C) and a bigger generation model (D) can't help if the answer was never retrieved.

  23. Q23D3 · Integration (incl. RAG)Select one

    recall@10 = 0.95 but answers include facts absent from the retrieved chunks. Which layer failed and what is the fix?

    • A. Retrieval; lower k and re-index.
    • B. Generation/grounding (hallucination); enforce answer-only-from-context, add citations, verify faithfulness, and rerank so the best passage is on top.
    • C. Embeddings; change the embedding model.
    • D. Indexing; re-index everything.
    Show answer

    Answer: B.

    High recall means the right context is present, so an unsupported fact is a grounding failure at generation. The retrieval-side fixes (A, C, D) target a healthy layer (recall-only reasoning).

  24. Q24D3 · Integration (incl. RAG)Select one

    A RAG system indexes with text-embedding-3-large but a new service queries with a different embedding model; similarity scores look random. What is the cause and fix?

    • A. The vector store is corrupt; rebuild the hardware.
    • B. Query/index embedding-model mismatch; use the identical embedding model for both indexing and querying.
    • C. k is too low; raise it to 1000.
    • D. The generation model is too small.
    Show answer

    Answer: B.

    Embeddings from different models occupy different vector spaces, so cross-model similarity is meaningless; query and index must use the same embedding model. It isn't hardware (A); raising k (C) can't fix incompatible vectors; the generation model (D) is unrelated to retrieval similarity.

  25. Q25D3 · Integration (incl. RAG)Select one

    A knowledge assistant keeps giving the OLD answer, confidently, after a document was updated overnight. What should the architect investigate FIRST?

    • A. Rewrite the system prompt to be more accurate.
    • B. Inspect what retrieval returned; the updated document was likely not re-chunked/re-embedded/re-indexed, so retrieval serves stale vectors.
    • C. Switch to Opus 5 with xhigh effort.
    • D. Add more few-shot examples.
    Show answer

    Answer: B.

    Confident-wrong immediately after a refresh points to stale retrieval/indexing; inspect the retrieved chunk IDs first. Prompt changes (A, D) and a bigger model (C) cannot fix stale retrieval.

  26. Q26D3 · Integration (incl. RAG)Select one

    A support agent has 22 tools, including delete_ticket and refund its role should never use. What is the correct remediation?

    • A. Keep the tools but log every call for audit.
    • B. Remove the unneeded tools from the agent's allowlist (least privilege) so it cannot invoke them at all.
    • C. Add an 'are you sure?' confirmation before those tools run.
    • D. Forbid them in the system prompt.
    Show answer

    Answer: B.

    Least privilege removes the capability so the excessive agency is gone. Logging (A) and confirmation (C) leave it reachable; a prompt rule (D) is prompt-as-enforcement and bypassable.

  27. Q27D3 · Integration (incl. RAG)Select two

    Contracts chunked at fixed 500 tokens produce answers citing the wrong sub-clause and losing surrounding context. Which TWO changes help MOST?

    • A. Structural/document-aware chunking that splits on clause/section boundaries.
    • B. Parent–child chunking: match small children, return the larger parent clause for context.
    • C. Dense-only retrieval.
    • D. Higher temperature.
    • E. Removing citations.
    Show answer

    Answer: A and B.

    Structured documents need boundary-aware chunking and parent–child context. Dense-only (C), temperature (D), and removing citations (E) don't address chunking and the last two make it worse.

  28. Q28D3 · Integration (incl. RAG)Select one

    A remote MCP server over Streamable HTTP exposes powerful tools with no authentication. What MUST be added?

    • A. Nothing; MCP is safe by default.
    • B. OAuth 2.1 on the remote MCP server plus per-user permission checks inside the tools.
    • C. A longer system prompt.
    • D. A higher rate-limit tier.
    Show answer

    Answer: B.

    Remote MCP requires OAuth 2.1 and per-user authorization so the agent acts with the caller's authority. MCP isn't authenticated by default (A); prompts (C) and rate limits (D) don't address authorization.

  29. Q29D3 · Integration (incl. RAG)Select one

    After enabling backoff retries, some customers are charged twice. What is the correct fix?

    • A. Disable retries entirely.
    • B. Add idempotency keys to the charge action so retried calls de-duplicate with no duplicate side effect.
    • C. Lower the model temperature.
    • D. Reconcile duplicate charges weekly.
    Show answer

    Answer: B.

    Idempotency keys make retries on non-idempotent actions safe. Disabling retries (A) harms resilience; temperature (C) is irrelevant; weekly reconciliation (D) still double-charged customers (silent failure).

  30. Q30D3 · Integration (incl. RAG)Select one

    A 3M-document corpus changes daily and answers must cite the exact clause. Which approach is BEST?

    • A. Fine-tune the model on the corpus every night.
    • B. RAG with hybrid retrieval, reranking, citations, and a scheduled refresh pipeline.
    • C. Stuff the whole corpus into the 1M context per query.
    • D. Long context plus fine-tuning combined.
    Show answer

    Answer: B.

    Large, daily-changing, citation-requiring corpora are canonical RAG. Nightly fine-tuning (A) is an infeasible cadence for facts; the corpus exceeds/overspends context and loses citations (C); (D) inherits both problems.

  31. Q31D3 · Integration (incl. RAG)Select two

    An agent retrieves k=8 with dense-only and no reranker; precision is poor and exact part numbers are missed. Which TWO changes give the biggest quality lift?

    • A. Switch to hybrid retrieval (BM25 + dense with rank fusion) so exact identifiers rank.
    • B. Retrieve wide (k≈50) and add a reranker, keeping the top 6.
    • C. Increase temperature.
    • D. Remove citations to speed responses.
    • E. Move to a 1M-token context and stuff everything.
    Show answer

    Answer: A and B.

    Hybrid retrieval fixes exact-identifier misses and retrieve-wide-then-rerank fixes precision. Temperature (C) is irrelevant to retrieval; removing citations (D) harms grounding traceability; stuffing context (E) inflates cost without improving ranking (over-engineered).

  32. Q32D4 · Evaluation, Testing and OptimizationSelect one

    On 2,000 paired items prompt B wins 1,080 (54%); on a separate 12-item hand test B won 7. Which result should drive promotion and why?

    • A. The 12-item test, because the answers looked clearly better.
    • B. The 2,000-item result: a 54% win at n=2,000 is ~3.6 standard deviations from chance (p < 0.001), significant — provided no segment regresses.
    • C. Neither; A/B testing is unreliable.
    • D. Average the two win rates.
    Show answer

    Answer: B.

    At n=2,000 the 80-win excess over 1,000 expected is ≈3.6σ (std dev ≈ 22.4), which is significant; 7/12 is coin-flip noise. The small test (A) is noise; A/B is reliable at scale (C); averaging win rates (D) is meaningless.

  33. Q33D4 · Evaluation, Testing and OptimizationSelect one

    A prompt change wins significantly overall but Spanish-tier accuracy drops 6 points. What is the correct decision?

    • A. Promote it; the overall win is significant.
    • B. Do not promote as-is; a per-segment regression blocks promotion even when the aggregate improves — fix the Spanish regression first.
    • C. Promote and monitor complaints.
    • D. Drop Spanish from the eval set.
    Show answer

    Answer: B.

    Per-segment no-regression is a hard gate; an aggregate win hiding a segment regression is anti-pattern #10. Promoting anyway (A, C) ships a known regression; dropping Spanish (D) re-hides it (aggregate-metric).

  34. Q34D4 · Evaluation, Testing and OptimizationSelect one

    A team can only report a single overall accuracy number and cannot tell which segment fails. Which observability gap is the root cause?

    • A. Missing GPU utilisation metrics.
    • B. Eval records lack a segment tag (and prompt/model version), so metrics can't be stratified or attributed.
    • C. The judge model is too small.
    • D. Latency is not logged.
    Show answer

    Answer: B.

    Without segment tags and version fields only an aggregate is computable and regressions can't be attributed. GPU metrics (A) are irrelevant; judge size (C) doesn't create the reporting gap; latency (D) is a different signal.

  35. Q35D4 · Evaluation, Testing and OptimizationSelect two

    Which TWO fields are MOST essential in a per-request eval record to support per-segment evaluation and A/B attribution?

    • A. A segment tag (tenant/language/type).
    • B. The prompt_version and model used.
    • C. The server's CPU temperature.
    • D. A random UUID with no linkage.
    • E. The marketing campaign name.
    Show answer

    Answer: A and B.

    Segment tags enable stratified metrics and prompt/model version enables attributing regressions and A/B results to a specific change. CPU temperature (C), an unlinked UUID (D), and a campaign name (E) don't support evaluation.

  36. Q36D4 · Evaluation, Testing and OptimizationSelect one

    An eval grades answers with the same Sonnet 5 session that produced them and reports high scores. Why is this unsound and what is the FIRST fix?

    • A. It costs too much; batch it.
    • B. Same-session self-review carries the producing bias; grade with a separate model/session and calibrate it against human labels.
    • C. The judge must always be a larger model.
    • D. LLM-as-judge is never valid.
    Show answer

    Answer: B.

    Grading in the producing session retains the original bias (anti-pattern #9); independence plus human calibration is required. Cost (A) isn't the core issue; the judge needn't be larger (C); LLM-as-judge is valid when independent (D).

  37. Q37D4 · Evaluation, Testing and OptimizationSelect one

    An interactive streaming UI feels slow even though total latency is acceptable. Which metric should be added?

    • A. Total tokens generated.
    • B. Time-to-first-token (TTFT), because perceived latency in streaming UIs is driven by how fast output starts.
    • C. Daily request count.
    • D. Number of prompt versions.
    Show answer

    Answer: B.

    In streaming UIs users perceive responsiveness by when tokens start (TTFT), not just total time. Total tokens (A), request count (C), and version count (D) are not latency-experience metrics.

  38. Q38D4 · Evaluation, Testing and OptimizationSelect one

    A system fails only on the hardest 5% of cases and is fine elsewhere. What is the most likely cause and fix?

    • A. Retrieval failure; re-chunk everything.
    • B. Model mismatch (tier too small for the hardest cases); escalate just those to a stronger model via a cascade.
    • C. Prompt failure; rewrite the whole prompt.
    • D. Infrastructure; add more replicas.
    Show answer

    Answer: B.

    A clean 'only the hardest cases fail' signature points to model capability; a cascade escalates just those cases while keeping the cheap model elsewhere. Re-chunking (A) and rewriting the prompt (C) target working layers; replicas (D) don't affect correctness.

  39. Q39D4 · Evaluation, Testing and OptimizationSelect one

    A team wants to cut cost 50% on a nightly bulk-summarisation job with no latency requirement. Which lever is BEST and what MUST follow?

    • A. Lower effort blindly and ship.
    • B. Move the job to the Message Batches API (50% discount, within 24 h) and re-run the per-segment regression suite to confirm no quality regression.
    • C. Switch interactive traffic to Haiku too.
    • D. Disable evaluation to save compute.
    Show answer

    Answer: B.

    A latency-tolerant bulk job fits Batch, and every cost change is followed by per-segment re-validation. Blind effort cuts (A) risk quality; changing interactive traffic (C) is out of scope; disabling evals (D) removes the safety net (silent failure).

  40. Q40D4 · Evaluation, Testing and OptimizationSelect two

    Which TWO are the right metrics for a high-volume, cost-sensitive service with an interactive SLA?

    • A. Cost per task.
    • B. p95 latency.
    • C. Mean latency only.
    • D. Total fleet tokens.
    • E. Prompt-version count.
    Show answer

    Answer: A and B.

    Cost-sensitive plus interactive points to cost per task and p95 tail latency. Mean latency (C) hides the tail; total tokens (D) and version count (E) are not user-facing SLA metrics (aggregate-metric).

  41. Q41D4 · Evaluation, Testing and OptimizationSelect one

    An LLM judge consistently rates the longer of two answers higher regardless of correctness. What is happening and the FIRST fix?

    • A. The judge is perfectly reliable; keep it.
    • B. Length bias; calibrate against human labels and revise the rubric to score correctness/grounding, not length.
    • C. Always prefer the longer answer.
    • D. Delete the eval entirely.
    Show answer

    Answer: B.

    The judge exhibits length bias, fixed by human calibration and a correctness-focused rubric. The judge is not reliable (A); rewarding length (C) is the bug; deleting evals (D) abandons measurement.

  42. Q42D5 · Governance, Safety and Risk ManagementSelect two

    A healthcare intake assistant handles PHI, is EU-based on AWS, and a vendor proposes Fable 5.1 with indefinite retention. Which TWO corrections are REQUIRED first?

    • A. Choose a ZDR-eligible model (Opus 5 / Sonnet 5) because Fable 5.1's 30-day retention fails a ZDR/PHI posture.
    • B. Deploy on Bedrock in an EU region and sign a BAA, with retention limits and a DSAR/erasure path.
    • C. Keep Fable 5.1 but promise to delete logs after 30 days.
    • D. Rely on a system-prompt rule 'never expose PHI'.
    • E. Retain all data indefinitely for quality.
    Show answer

    Answer: A and B.

    A PHI/ZDR posture disqualifies Fable 5.1 and requires EU residency, a BAA, retention limits and erasure. 30-day deletion (C) is exactly what ZDR forbids; a prompt rule (D) can't protect PHI (prompt-as-enforcement); indefinite retention (E) breaches GDPR minimisation.

  43. Q43D5 · Governance, Safety and Risk ManagementSelect one

    A US federal agency mandates FedRAMP High and wants the newest model the day it ships. What is the correct guidance?

    • A. Use the direct Anthropic API for earliest access.
    • B. Deploy Claude via Bedrock or Vertex inside the FedRAMP High boundary; the authorisation requirement overrides the desire for day-one access.
    • C. Any cloud works; Claude is inherently authorised.
    • D. Run it on a laptop in the office.
    Show answer

    Answer: B.

    FedRAMP High is provided through the Bedrock/Vertex boundary and the compliance requirement is binding over novelty. Direct API (A) is outside the boundary (constraint-blind); authorisation isn't inherent (C); a laptop (D) is unacceptable for federal data.

  44. Q44D5 · Governance, Safety and Risk ManagementSelect one

    An incident-response runbook for an EU personal-data system is missing one legally required element. Which is it?

    • A. A marketing statement.
    • B. The GDPR breach-notification step and 72-hour timeline to the supervisory authority.
    • C. A list of favourite dashboards only.
    • D. A cost projection.
    Show answer

    Answer: B.

    GDPR requires breach notification (generally within 72 hours), so the runbook must include it. Marketing (A), a dashboard list (C), and a cost projection (D) are not the legal requirement.

  45. Q45D5 · Governance, Safety and Risk ManagementSelect two

    An agent summarising web pages reads hidden text instructing it to email internal data externally. What is it and the correct defence?

    • A. Indirect prompt injection.
    • B. Treat external content as untrusted, apply least privilege so it lacks an unrestricted email tool, and validate/deny such actions.
    • C. Trust it because it came through a legitimate tool.
    • D. Increase the model's effort level.
    • E. Add the instruction to the system prompt.
    Show answer

    Answer: A and B.

    Malicious instructions in fetched content are indirect prompt injection; defence is untrusted-content handling plus least privilege and action validation. Trusting tool output (C) is the vulnerability; effort (D) is irrelevant; adding it to the prompt (E) executes the attack.

  46. Q46D5 · Governance, Safety and Risk ManagementSelect one

    A support system escalates to a human whenever the customer 'sounds angry'. Why is this flawed?

    • A. It is optimal.
    • B. Sentiment is not a proxy for complexity or risk; escalate on measured complexity/risk signals or deterministic rules instead.
    • C. It escalates too rarely.
    • D. Anger always means the model failed.
    Show answer

    Answer: B.

    Sentiment-based escalation (anti-pattern #5) conflates emotion with difficulty; a calm customer can have a complex, high-risk case. It is not optimal (A); frequency (C) isn't the core flaw; anger doesn't imply model failure (D).

  47. Q47D5 · Governance, Safety and Risk ManagementSelect one

    A stem states 'high-risk processing of EU residents' personal data'. Which control belongs at the DESIGN gate?

    • A. Wait for a complaint before assessing.
    • B. Complete a DPIA and decide EU residency/placement so residency, retention and erasure shape the architecture before build.
    • C. Add a UI disclaimer at launch.
    • D. Store everything to be safe.
    Show answer

    Answer: B.

    GDPR high-risk processing requires a DPIA and residency/erasure decisions at the design gate. Waiting (A) is too late; a disclaimer (C) doesn't satisfy the DPIA; over-retention (D) breaches minimisation.

  48. Q48D5 · Governance, Safety and Risk ManagementSelect one

    Which set best represents layered guardrails (defence in depth)?

    • A. A single, very detailed system prompt.
    • B. Input classifier → system-prompt guidance → tool-permission hooks → output validation → human review for high-stakes actions.
    • C. Only human review at the end.
    • D. Only an output regex.
    Show answer

    Answer: B.

    Defence in depth stacks independent layers so a bypass of one is caught by the next. A single prompt (A), review-only (C), or regex-only (D) are single points of failure.

  49. Q49D5 · Governance, Safety and Risk ManagementSelect one

    A prompt-injection incident is detected in production. What is a sound FIRST containment step?

    • A. Delete all logs.
    • B. Roll back the prompt/model version and/or disable the exploited tool via its permission hook, scoping impact with correlation IDs and traces.
    • C. Ignore it until the next release.
    • D. Publicly post customer data for transparency.
    Show answer

    Answer: B.

    Containment is fast rollback and disabling the exploited capability, scoped via traces. Deleting logs (A) destroys evidence; ignoring (C) prolongs harm; disclosing data (D) is a second breach.

  50. Q50D5 · Governance, Safety and Risk ManagementSelect two

    Which TWO signal-word-to-control mappings are correct?

    • A. 'FedRAMP High' → deploy via Bedrock/Vertex inside the authorised boundary.
    • B. 'ZDR mandated' → do not use Fable 5.1 (30-day retention); use a ZDR-eligible model.
    • C. 'PHI' → a strongly worded system prompt is sufficient.
    • D. 'EU erasure' → retain data indefinitely.
    • E. 'high-risk AI' → skip documentation.
    Show answer

    Answer: A and B.

    FedRAMP maps to the cloud boundary and a ZDR mandate disqualifies Fable 5.1. PHI needs a BAA and controls, not a prompt (C); EU erasure requires a deletion path, not indefinite retention (D); high-risk AI (EU AI Act) requires documentation, not skipping it (E).

  51. Q51D5 · Governance, Safety and Risk ManagementSelect one

    Where should API keys and DB credentials for a Claude agent live?

    • A. In the system prompt so the model can use them.
    • B. In environment variables or a secret manager, never in prompts, CLAUDE.md, or logs.
    • C. In CLAUDE.md for convenience.
    • D. Hard-coded in tool source and printed to logs.
    Show answer

    Answer: B.

    Secrets belong in env/secret managers and never in model-visible config or logs. Putting them in the prompt (A), CLAUDE.md (C), or logs (D) creates exfiltration paths.

  52. Q52D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    A VP sponsor was shown a 40-page technical ADR and remains unconvinced. What is the BEST corrective communication?

    • A. Resend the ADR with a summary email.
    • B. Present a one-page decision matrix with weighted criteria, a cost model, and business value/risk framed for an executive.
    • C. Send the raw evaluation logs.
    • D. Give the VP the operator runbook.
    Show answer

    Answer: B.

    Sponsors need a decision matrix/cost model/one-page brief at their altitude, not an engineering ADR. Resending the ADR (A) repeats the wrong-altitude error; raw eval logs (C) and the runbook (D) are the wrong artefacts for a sponsor.

  53. Q53D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    Using a weighted decision matrix (SLA 0.25, cost 0.25, reliability 0.20, flexibility 0.15, effort 0.15), Workflow scores 5,5,5,3,4 and Agentic scores 3,3,3,5,3. Which wins and why?

    • A. Agentic, because it is more flexible.
    • B. Workflow (4.60 vs 3.30), because the business weighted SLA, cost and reliability highest.
    • C. They tie.
    • D. The matrix cannot decide.
    Show answer

    Answer: B.

    Workflow = 0.25·5+0.25·5+0.20·5+0.15·3+0.15·4 = 4.60; Agentic = 0.25·3+0.25·3+0.20·3+0.15·5+0.15·3 = 3.30, so workflow wins on the business-weighted criteria. Flexibility (A) is only 0.15; it isn't a tie (C); the matrix decides transparently (D).

  54. Q54D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    A sponsor asks for 'faster support'. How should this become an SLA?

    • A. Promise it will be 'much faster'.
    • B. Agree numeric per-segment targets (p95 latency, deflection, cost/ticket) and report against them on a cadence.
    • C. Track mean latency only.
    • D. Leave it undefined and adjust later.
    Show answer

    Answer: B.

    SLAs must be measurable and per-segment, mirroring eval criteria, with regular reporting. Vague promises (A) can't be verified; mean-only (C) hides the tail; undefined (D) invites disputes.

  55. Q55D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    Ops refuses a handoff. Which deliverable set satisfies the handoff exit criterion?

    • A. A slide deck and a promise of support.
    • B. A runbook (dashboards, alerts, on-call, rollback, dependencies) plus C4 and data-flow docs, so operators can run and recover the system.
    • C. The source code repository only.
    • D. A cost model.
    Show answer

    Answer: B.

    Handoff's exit criterion is operability and recoverability, provided by the runbook plus architecture/data-flow docs. A deck (A), code alone (C), or a cost model (D) don't let operators run and recover it.

  56. Q56D6 · Stakeholder Communication and Lifecycle ManagementSelect two

    A model a service pins was retired. Which TWO steps make the migration correct?

    • A. Read the deprecation notice, identify breaking changes, and re-run the regression suite per segment on the target model, adjusting prompts.
    • B. Canary the new model with rollback, then ramp, and communicate timeline/differences to stakeholders.
    • C. Swap the model ID in production and hope for the best.
    • D. Keep calling the retired model.
    • E. Disable evals during migration to move faster.
    Show answer

    Answer: A and B.

    Deprecation is a governed lifecycle event: assess breaking changes, re-validate per segment, canary with rollback, and communicate. A blind swap (C) risks regressions; calling a retired model (D) fails; disabling evals (E) removes the safety net.

  57. Q57D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    During discovery a stakeholder says 'make onboarding faster with AI' and offers no numbers. What is the FIRST step?

    • A. Choose Opus 5 and start building.
    • B. Run the success-metric step to turn 'faster' into numeric, per-segment criteria, capture constraints, and get sign-off before design.
    • C. Assume a 50% improvement target.
    • D. Escalate to the vendor.
    Show answer

    Answer: B.

    Discovery converts vague asks into numeric, per-segment criteria with constraints and sign-off before design. Building (A), assuming a target (C), or escalating (D) skip the anchoring step (constraint-blind).

  58. Q58D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    What is the correct distinction among SLA, SLO and SLI?

    • A. They are synonyms.
    • B. SLI is the measured indicator, SLO is the internal target, SLA is the external commitment (often with consequences).
    • C. SLA is internal; SLO is external.
    • D. SLI is a legal contract.
    Show answer

    Answer: B.

    SLI (indicator) → SLO (internal objective) → SLA (external commitment). They are not synonyms (A), the internal/external roles are not swapped (C), and the SLI is a metric, not a contract (D).

  59. Q59D6 · Stakeholder Communication and Lifecycle ManagementSelect two

    A technically strong system sees low adoption. Which TWO change-management levers help MOST?

    • A. Enablement/training covering capabilities and limits, and how to review AI output.
    • B. A phased rollout with feedback capture and value reporting against agreed metrics.
    • C. Mandating org-wide use immediately with no support.
    • D. Removing human-in-the-loop gates to feel faster.
    • E. Hiding limitations from users.
    Show answer

    Answer: A and B.

    Adoption is driven by enablement and phased rollout with feedback and demonstrated value. Forced use (C), removing gates (D), and hiding limits (E) backfire.

  60. Q60D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    Which artefact best communicates residual risk (after mitigation) to a compliance stakeholder?

    • A. An ADR listing the technical decision only.
    • B. A risk register scoring each risk by likelihood × impact with owner, mitigation layer, and remaining residual risk.
    • C. A cost model showing unit economics.
    • D. A C4 code diagram.
    Show answer

    Answer: B.

    A risk register with likelihood/impact, owners, mitigations and residual risk is the compliance-facing artefact. An ADR (A) records a technical decision; a cost model (C) is economics; a code diagram (D) is for engineers.

  61. Q61D7 · Developer Productivity and Operational EnablementSelect one

    A platform team must guarantee git push --force is blocked for every developer and unbypassable. Where does this belong?

    • A. A sentence in ./CLAUDE.md.
    • B. A PreToolUse hook (exit code 2 blocks the call) delivered via managed policy so it cannot be overridden.
    • C. A note in each developer's CLAUDE.local.md.
    • D. A Slack reminder.
    Show answer

    Answer: B.

    Deterministic, non-bypassable blocking is a PreToolUse hook (exit 2) enforced through managed policy. A CLAUDE.md sentence (A) is prose guidance (prompt-as-enforcement); CLAUDE.local.md (C) is personal/overridable; Slack (D) is not enforcement.

  62. Q62D7 · Developer Productivity and Operational EnablementSelect one

    A team wants Claude to review diffs automatically in CI. What is the correct setup?

    • A. Run interactive Claude and have a human paste the diff.
    • B. Use headless mode: claude -p '…' --output-format json with a restricted --allowedTools and an appropriate --permission-mode, so CI can parse findings and gate the build.
    • C. Give the CI agent all tools for flexibility.
    • D. Disable permissions in CI to avoid friction.
    Show answer

    Answer: B.

    CI is non-interactive, so headless mode with structured output and restricted tools is correct. Interactive use (A) can't run in CI; all-tools (C) and disabled permissions (D) violate least privilege.

  63. Q63D7 · Developer Productivity and Operational EnablementSelect two

    A platform group is rolling Claude Code to 15 teams and wants low risk. Which TWO choices are BEST?

    • A. Pilot with 2 teams and telemetry, then champions to spread practice, then scale org-wide with outcome/cost dashboards.
    • B. Enforce the deny-list and required review hook via managed policy (non-overridable) plus a checked-in .claude/ catalogue.
    • C. Mandate org-wide use on day one to maximise ROI.
    • D. Let each team choose its own unvetted MCP servers.
    • E. Measure success by lines of code generated.
    Show answer

    Answer: A and B.

    Phased rollout plus managed-policy enforcement and a shared checked-in catalogue is the safe pattern. A day-one mandate (C) skips trust-building; unvetted MCP servers (D) break governance; lines of code (E) is a vanity metric.

Last updated Sep 18, 2026