CCAR-P Practice Exam
A full-length, 63-item timed practice exam for the Claude Certified Architect – Professional credential, distributed by blueprint weight with answer key and score interpretation.
This is a full-length, 63-item practice exam for CCAR-P. The questions are new (not reused from the domain pages) and distributed to match the blueprint. Sit it under exam conditions before you rely on the score.
Instructions
- Time: 120 minutes. Set a timer and do not pause it.
- Format: multiple-choice (select one) and multiple-response (select two); each stem states how many to select.
- Scoring: the real exam is scaled 100–1000 with a pass at 720. There is no guessing penalty — answer everything. As a raw proxy, aim for ≥ 80% (≈ 50/63) before booking.
- Method: read the whole stem, identify the binding constraint, eliminate the constraint-blind / over-engineered / prompt-as-enforcement / self-report / silent-failure / recall-only / aggregate-metric distractors, then choose.
Domain distribution
| # | Domain | Weight | Items here |
|---|---|---|---|
| 1 | Solution Design & Architecture | 17% | 11 |
| 2 | Claude Models, Prompting & Context Engineering | 13% | 8 |
| 3 | Integration (incl. RAG) | 19% | 12 |
| 4 | Evaluation, Testing & Optimization | 16% | 10 |
| 5 | Governance, Safety & Risk Management | 14% | 9 |
| 6 | Stakeholder Communication & Lifecycle Management | 14% | 9 |
| 7 | Developer Productivity & Operational Enablement | 7% | 4 |
| Total | 100% | 63 |
Score interpretation
| Raw score (of 63) | Percent | Reading |
|---|---|---|
| 57–63 | 90–100% | Exam-ready; strong across all domains |
| 50–56 | 80–89% | Likely pass; shore up your weakest domain |
| 45–49 | 71–79% | Borderline; drill D3/D1/D4 and re-sit |
| 38–44 | 60–70% | Not ready; systematic gaps — restudy by domain |
| < 38 | < 60% | Restudy the full course before re-attempting |
The real exam scales 100–1000 with a pass at 720; treat ≥ 80% raw (≈ 50/63) here as your go/no-go line, and confirm you have no single domain far below the rest.
Take the practice exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
63 questions · one at a time · 120-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A logistics firm wants to automate shipment-status replies: 500/hour, one internal API, p95 under 4 s, tight cost ceiling. Which pattern should the architect propose?
Show answer
Answer: B.
Known steps plus binding latency/cost favour the simplest pattern. Multi-agent (A) and broad agentic loops (C) over-engineer; fine-tuning (D) is premature.
A German bank requires all data to stay in the EU and already runs on AWS. Which placement is BEST?
Show answer
Answer: B.
EU residency plus an existing AWS footprint point to Bedrock EU. Direct API (A) risks residency; Vertex (C) ignores the AWS investment; novelty (D) can't override residency.
An agentic loop occasionally runs forever. Which is the CORRECT primary control plus backstop?
Show answer
Answer: B.
Control flow keys off
stop_reason; a cap is a backstop, not the mechanism. Text parsing (A) is brittle; a cap alone (C) is an anti-pattern; temperature (D) doesn't govern termination.Peak load is 25 req/s at ~10k input tokens each, and the team hits ITPM limits. Which TWO are sound?
Show answer
Answer: A and C.
Batch for tolerant work plus backoff, and raising the tier / spillover routing, are the correct capacity levers. Tight-loop retry (B) worsens it; swallowing 429s (D) is silent failure; Opus for all (E) raises cost without fixing limits.
What makes an end-to-end reference architecture complete at Professional level?
Show answer
Answer: B.
The feedback loop is mandatory; a design with no path from production signals to improvement is incomplete. The others omit it.
A stem gives a hard $8,000/month budget and a modest accuracy bar for high-volume classification. Which posture is BEST?
Show answer
Answer: B.
The binding pillar is cost at volume with a modest quality need — the cheapest adequate model with caching. Opus/Fable (A, C) overspend; multi-agent (D) adds cost.
When is multi-agent orchestration genuinely justified?
Show answer
Answer: B.
Parallel, context-isolated subtasks with fan-in are the multi-agent signal. Sequential shared work (A) is a workflow; a single lookup (C) is augmented-LLM; 'always' (D) is over-engineering.
The primary model returns sustained 529s under a spike. Which fallback is BEST?
Show answer
Answer: B.
Retry-then-fallback with degraded-mode logging preserves availability and observability. Hard failure (A) is poor HA; stale-as-fresh (C) is silent failure; downgrading a Fable 5.1 session (D) drops thinking blocks.
The company lacks ops capacity, needs launch in a month, and the capability is commodity. Build or buy?
Show answer
Answer: B.
No ops capacity + tight deadline + commodity capability → managed/buy. Building (A) contradicts the constraints; delaying (C) misses the deadline; (D) abandons the requirement.
Why capture a design as an ADR?
Show answer
Answer: B.
ADRs make reasoning and trade-offs explicit and reviewable. They are not a formality (A), don't replace evals (C), and aren't a technical prerequisite (D).
A stakeholder gives a vague goal and no numbers. What is the BEST first action?
Show answer
Answer: B.
Without criteria/constraints the design is unanchored; discovery surfaces the binding constraint. Building (A), defaulting to a model (C), or inventing a target (D) skip anchoring.
A pipeline runs everything on Opus 5 at 4× budget with acceptable quality. Best cost fix preserving quality?
Show answer
Answer: B.
Cheap-first with escalation on validation failure preserves hard-case quality; caching amortises the prefix. Blanket Haiku (A) sacrifices quality; self-report (C) is an anti-pattern; xhigh everywhere (D) raises cost.
After upgrading to Sonnet 5, requests setting
budget_tokensfail with 400. What is the fix?Show answer
Answer: B.
budget_tokensis removed on Sonnet 5 (returns 400); adaptive thinking + effort replaces it. Retries (A) can't fix a 400; Haiku (C) is not equivalent; removing thinking (D) discards needed reasoning.A cache hit rate is near zero because the volatile user turn is placed before the system prompt and documents. What is the fix?
Show answer
Answer: B.
Caching keys on a stable prefix; the volatile turn must come last. Disabling (A) forfeits savings; length (C) only matters for the minimum; model choice (D) is irrelevant.
In an active Fable 5.1 session the team must change the system instructions mid-run. Correct approach?
Show answer
Answer: B.
Fable 5.1 thinking blocks are invalidated by editing/reordering/removing earlier turns, so mid-session changes are appended. The other options invalidate downstream thinking.
A hard rule 'no discount above 25%' must be guaranteed. Where does it belong?
Show answer
Answer: B.
Hard rules require deterministic enforcement. Prompt wording (A) and few-shot (C) are prompt-as-enforcement; the thinking budget (D) is unrelated.
A workload mandates Zero Data Retention but the team wants Fable 5.1. Correct guidance?
Show answer
Answer: B.
Fable 5.1's 30-day retention disqualifies it under a ZDR mandate. ZDR isn't universal (A); your logging (C) doesn't change provider retention; 30-day deletion (D) is what ZDR forbids.
Which statement about current-model thinking config is correct?
Show answer
Answer: B.
Adaptive thinking + effort is standard; Haiku 4.5 is the
budget_tokensexception withouteffort. The others misstate the API.Reusable capability blocks are pasted into every prompt, bloating context and breaking the cache prefix. Better pattern?
Show answer
Answer: A.
Skills load capability on demand and keep the prefix stable. Duplication (B) caused the bloat; the volatile turn (C) worsens caching; a bigger window (D) doesn't fix cost/caching.
A knowledge assistant keeps giving the OLD answer, confidently, after a document was updated overnight. What to investigate FIRST?
Show answer
Answer: B.
Confident-wrong-after-refresh points to stale retrieval/indexing; inspect retrieved chunks first. Prompt changes (A, D) and a bigger model (C) can't fix stale retrieval.
A support agent has 18 tools including
delete_accountit should never use. Correct remediation?Show answer
Answer: B.
Least privilege removes the capability. Logging (A) and confirmation (C) leave it reachable; a prompt rule (D) is bypassable.
Users search by exact SKU and by description; dense-only retrieval misses many SKUs. Best fix?
Show answer
Answer: B.
Dense is weak on exact tokens; sparse handles SKUs and hybrid covers both. Huge k (A) adds noise; a bigger model (C) doesn't change retrieval; removing filters (D) harms precision/security.
recall@10 is 0.94 but answers include facts not in the retrieved chunks. Which layer and fix?
Show answer
Answer: B.
High recall means retrieval works; unsupported facts are a grounding failure. The retrieval-side fixes (A, C, D) target a healthy layer.
A 3M-document corpus changes daily and answers must cite the exact clause. Best approach?
Show answer
Answer: B.
Large, changing, citation-requiring corpora are canonical RAG. Fine-tuning (A) can't track daily facts; the corpus exceeds/overspends context and loses citations (C); (D) inherits both issues.
Contracts chunked at fixed 500 tokens produce answers citing the wrong sub-clause and losing context. Which TWO help most?
Show answer
Answer: A and B.
Structured docs need boundary-aware chunking and parent–child context. Dense-only (C), temperature (D), and removing citations (E) don't address chunking.
A remote MCP server over Streamable HTTP exposes powerful tools with no auth. What must be added?
Show answer
Answer: B.
Remote MCP requires OAuth 2.1 and per-user authorization. MCP isn't authenticated by default (A); prompts (C) and tiers (D) don't address authz.
After enabling backoff retries, some customers are charged twice. Correct fix?
Show answer
Answer: B.
Idempotency keys make retries safe. Disabling retries (A) harms resilience; temperature (C) is irrelevant; reconciling later (D) still double-charged.
Some questions require follow-up lookups combining multiple sources. Which design fits and its trade-off?
Show answer
Answer: A.
Multi-hop queries justify agentic RAG; the trade-off is cost/latency. Stuffing (B) doesn't scale; fine-tuning (C) can't hold changing facts; removing reranking (D) hurts precision.
A capability will be reused across many agents/clients under a standard protocol. Best integration mechanism?
Show answer
Answer: B.
Cross-client reuse under a standard protocol is MCP's purpose. Per-app tools (A) fragment; agent-to-agent (C) is for context isolation; prompt-pasting (D) is unmaintainable.
An agent loads all 40 tools and 30 docs each request: high cost, low cache hits, poor tool selection. Which TWO fix it?
Show answer
Answer: A and B.
Progressive discovery keeps context lean, restores the cache prefix, and improves tool selection. A bigger window (C) still pays for bloat; escalating (D) raises cost; disabling caching (E) is the opposite of the fix.
Which observability signals let you reconstruct a failed multi-step request end to end?
Show answer
Answer: A and B.
Correlation IDs plus per-span traces (with retrieval and cost detail) make an incident reconstructable. A status code (C) and daily counts (D) are too coarse; self-assessment (E) is unreliable.
A bot reports 90% overall accuracy but only refund tickets fail. What is needed?
Show answer
Answer: B.
Aggregate masks a segment failure; stratified metrics expose it. The aggregate (A) misleads; a same-mix set (C) still hides it; blanket escalation (D) overpays.
Grading answers with the same model and session that produced them is unsound because…
Show answer
Answer: B.
Independence plus human calibration is required. Cost (A) isn't the core issue; the judge needn't be larger (C); LLM-as-judge is valid when independent (D).
Prompt B beats A on 7 of 12 hand-picked examples. Correct conclusion?
Show answer
Answer: B.
7/12 is noise; promotion needs a significant paired result with no per-segment regression. Eyeballing (A, C) and averaging prompts (D) are unsound.
Mean latency is 2 s but users complain; p95 is 12 s. What should the SLA track?
Show answer
Answer: B.
Interactive SLAs track p95/p99. The mean (A) hides the tail; count (C) and tokens (D) aren't latency SLAs.
An LLM judge rates longer answers higher regardless of correctness. What is it and the fix?
Show answer
Answer: A and B.
Length bias is fixed by calibration and a correctness-focused rubric. The judge isn't reliable (C), rewarding length (D) is the bug, and deleting evals (E) abandons measurement.
How to prevent a fix for one segment from silently breaking another?
Show answer
Answer: B.
A per-segment regression gate in CI is the guard. Spot-checks (A) and author trust (C) are unreliable; testing only the changed segment (D) is the risk.
A system fails only on the hardest 5% of cases. Most likely cause and fix?
Show answer
Answer: B.
'Only the hardest cases fail' signals model capability; a cascade escalates those cases. Re-chunking (A) and rewriting the prompt (C) target working layers; replicas (D) don't affect correctness.
Cut cost 50% on a nightly bulk job with no latency need. Best lever and what must follow?
Show answer
Answer: B.
Latency-tolerant bulk work fits Batch; every cost change is followed by per-segment re-validation. Blind effort cuts (A) risk quality; changing interactive traffic (C) is out of scope; disabling evals (D) removes the safety net.
Which two are the right metrics for a high-volume, cost-sensitive service with an interactive SLA?
Show answer
Answer: A and B.
Cost-sensitive + interactive → cost per task and p95. Mean (C) hides the tail; total tokens (D) and version count (E) aren't user-facing SLA metrics.
A golden set is 85% English tickets; production fails on Spanish. Flaw and fix?
Show answer
Answer: B.
An unstratified set makes a segment invisible offline. Size (A) isn't the issue; the model isn't broken (C); online evals (D) are still needed but the root cause is coverage.
A payment agent must never move over $10,000 without approval. Correct design?
Show answer
Answer: B.
Hard limits need deterministic enforcement plus a human gate. Prompt wording (A), self-report (C), and sentiment (D) are anti-patterns.
A US federal agency requires FedRAMP High. Compliant deployment?
Show answer
Answer: B.
FedRAMP High is provided through Bedrock/Vertex boundaries. Direct API (A) is outside it; compliance isn't inherent (C); a laptop (D) is unacceptable.
An agent summarising web pages reads hidden text telling it to email data externally. What is it and the defence?
Show answer
Answer: A and B.
Malicious instructions in fetched content are indirect injection; defence is untrusted-content handling plus least privilege and validation. Trusting tool output (C) is the vulnerability; effort (D) is irrelevant; adding it to the prompt (E) executes the attack.
A clinic wants a Claude assistant reading patient records. First compliance requirements?
Show answer
Answer: A and B.
PHI under HIPAA requires a BAA and PHI controls with auditing. 'Internal' (C) doesn't exempt PHI; publishing data (D) is a violation; sentiment escalation (E) is unrelated.
Escalating to a human whenever the customer 'sounds angry' is flawed because…
Show answer
Answer: B.
Sentiment-based escalation conflates emotion with difficulty. It isn't optimal (A); frequency (C) isn't the core flaw; anger doesn't imply model failure (D).
Which set best represents layered guardrails?
Show answer
Answer: B.
Defence in depth stacks independent layers. A single prompt (A), review-only (C), or regex-only (D) are single points of failure.
High-risk processing of EU residents' personal data — required before launch?
Show answer
Answer: A and B.
GDPR high-risk processing requires a DPIA plus residency and DSAR/erasure support. Waiting (C) and indefinite storage (D) violate GDPR; confidence escalation (E) is unrelated.
A prompt-injection incident is detected in production. Sound first containment?
Show answer
Answer: B.
Containment is fast rollback and disabling the exploited capability, scoped via traces. Deleting logs (A) destroys evidence; ignoring (C) prolongs harm; disclosing data (D) is a second breach.
Where should API keys and DB credentials for a Claude agent live?
Show answer
Answer: B.
Secrets belong in env/secret managers, never in model-visible config or logs. The others create exfiltration paths.
A hiring-screening assistant makes consequential recommendations. Required controls?
Show answer
Answer: A and B.
Consequential people-decisions require fairness evaluation and human authority. Full automation (C) launders bias; sentiment escalation (D) and indefinite retention (E) are anti-patterns/GDPR issues.
Two stakeholders disagree and no success metric exists; the team wants to build. First action?
Show answer
Answer: B.
Without agreement or metrics the discovery gate isn't passed. Building (A), picking a model (C), or escalating (D) skip anchoring.
Presenting a workflow-vs-agentic decision to an executive sponsor AND engineers. Best approach?
Show answer
Answer: B.
Communication matches audience altitude. One document for both (A, C) or raw logs (D) mismatches an audience.
Turning 'faster support' into an SLA. Correct approach?
Show answer
Answer: B.
SLAs must be measurable and per-segment with reporting. Vague promises (A) can't be verified; mean-only (C) hides the tail; undefined (D) invites disputes.
Which C4 level best shows executives how the system fits its environment?
Show answer
Answer: B.
The Context level shows the system, users, and external systems for a broad audience. Code (A) and Component (C) are for engineers; exhaustive diagrams (D) overwhelm.
A model the service pins was retired. Correct migration approach?
Show answer
Answer: A and B.
Deprecation is a governed lifecycle event: assess, re-validate per segment, canary, communicate. A blind swap (C) risks regressions; calling a retired model (D) fails; disabling evals (E) removes the safety net.
What deliverable is required to pass the handoff gate to operations?
Show answer
Answer: B.
Handoff's exit criterion is that operators can run and recover the system. A deck (A), nothing (C), or code alone (D) don't enable safe operation.
When should a DPIA/BAA be addressed in the lifecycle?
Show answer
Answer: B.
Compliance must shape design, so DPIA/BAA belong to the design exit criterion. Later (A, D) is too late; 'internal' (C) doesn't exempt regulated data.
A technically strong system sees low adoption. Which levers help most?
Show answer
Answer: A and B.
Adoption is driven by enablement, phased rollout, feedback, and demonstrated value. Forced use (C), removing gates (D), and hiding limits (E) backfire.
The team plans to deploy to 100% of users with no runbook and no agreed metrics. What does the correct answer insert?
Show answer
Answer: A and B.
The missing gates are agreed criteria and a runbook plus canary with rollback. Full deployment (C) is high blast radius; skipping monitoring (D) blinds ops; per-engineer metrics (E) destroy a shared definition of success.
A platform team wants a coding standard and denied destructive commands enforced across all repos with no override. Best mechanism?
Show answer
Answer: B.
Non-overridable, org-wide enforcement is what managed policy provides, with checked-in project config as the baseline.
CLAUDE.local.md(A) is git-ignored/overridable; Slack (C) and email (D) aren't enforcement.Setting up automated diff review in CI. Correct approach?
Show answer
Answer: B.
CI is non-interactive → headless with structured output and restricted tools. Interactive use (A) can't run in CI; all-tools (C) and disabled permissions (D) break least privilege.
A manager wants to measure AI productivity by lines of code generated and suggestions accepted. Guidance?
Show answer
Answer: B.
Productivity ties to delivery outcomes, not volume. Lines/acceptances (A) and tokens (C) are vanity signals; not measuring (D) forfeits value demonstration.
Last updated Sep 18, 2026