CCAR-P Practice Exam 2
A second full-length, blueprint-weighted CCAR-P practice exam with new, harder, multi-constraint items.
This is a second full-length, 63-item practice exam for CCAR-P. Every item is new — none appears in Practice Exam 1 or on the domain pages — and the set is tuned to be slightly harder than exam 1: more multi-constraint stems, more capacity/cost arithmetic, and more FIRST / MOST cost-effective / TWO qualifiers that reward reading the whole stem before answering. The blueprint weighting is identical to exam 1, so your per-domain scores are directly comparable.
What is different from exam 1
- Harder qualifiers. Several stems ask for the
FIRSTaction or theMOST cost-effectivedesign, so more than one option is defensible and you must pick the best. - Arithmetic items. You will compute per-request cost, RPM/ITPM/OTPM, cache savings, recall@k, MRR, and A/B significance from the numbers given.
- Multi-constraint scenarios. Residency + existing cloud + irreversible action, or PHI + ZDR + placement, must all be satisfied at once.
- Same distractor families as exam 1 (constraint-blind, over-engineered, prompt-as-enforcement, self-report, silent failure, recall-only, aggregate-metric, sentiment-as-complexity) so practising exam 1 transfers directly.
Domain distribution
| # | Domain | Weight | Items here |
|---|---|---|---|
| 1 | Solution Design & Architecture | 17% | 11 |
| 2 | Claude Models, Prompting & Context Engineering | 13% | 8 |
| 3 | Integration (incl. RAG) | 19% | 12 |
| 4 | Evaluation, Testing & Optimization | 16% | 10 |
| 5 | Governance, Safety & Risk Management | 14% | 9 |
| 6 | Stakeholder Communication & Lifecycle Management | 14% | 9 |
| 7 | Developer Productivity & Operational Enablement | 7% | 4 |
| Total | 100% | 63 |
How to use both exams
-
Sit Exam 1 first, under timed conditions, before you touch exam 2. Use it to find your weak domains and restudy those domain pages.
-
Use Exam 2 as the booking gate. Only after Exam 1 is comfortable (≥ 80% raw, no single domain far below the rest) should you sit this harder set. Treat a strong Exam 2 as the signal that you are ready to book.
-
Compare per-domain results across both sittings. A domain that drops sharply from exam 1 to exam 2 is where the deeper, multi-constraint reasoning is still shaky — return to that domain page’s scenario walkthrough and misconceptions table.
Scoring proxy
The real exam scales 100–1000 with a pass at 720; there is no guessing penalty, so answer everything. As a raw proxy, aim for ≥ 80% (≈ 50/63) on this harder set before booking, and confirm no single domain is far below the others.
Take the practice exam
Interactive mode
Take the practice exam
63 questions · one at a time · 120-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A workload runs 10 req/s average at 5k input + 1k output tokens on Sonnet 5 (
$2/$10per MTok) and must fit a$300k/month budget. Which design is the MOST cost-effective while preserving quality?Show answer
Answer: B.
Per-request cost is
5k×$2/1e6 + 1k×$10/1e6 = $0.02; cost-effectiveness is an arithmetic exercise of cache + cascade + batch layered until cost < budget. Blanket Opus (A) is over-engineered/overspends; guessing (C) is unanchored (constraint-blind); Fable 5.1 (D) is the most expensive model, the opposite of cost-effective.Peak load is 25 req/s at 8k input tokens each; the tier caps ITPM at 6,000,000. What is the binding constraint and the FIRST sound response?
Show answer
Answer: B.
peak_ITPM = 25×60×8,000 = 12,000,000, double the 6M cap, so ITPM binds first; the fix is Batch, a higher tier, or spillover. RPM-only (A) misreads the maths; Opus for all (C) raises cost and ITPM; ignoring bursts (D) invites a 429 storm (silent-failure/constraint-blind).
A p95 SLA of 5 s is being exceeded; traces show a long generated output dominates the latency budget. Which change is BEST?
Show answer
Answer: B.
When generation dominates the latency budget, shrink model/output/effort or stream. Multi-agent (A) is over-engineered and adds latency; larger k (C) adds retrieval time; Batch (D) has up-to-24 h latency and cannot serve an interactive p95 SLA.
An EU-regulated insurer on Azure wants claims triaged with a fraud flag that must NEVER auto-deny a claim. Which TWO design choices are correct?
Show answer
Answer: A and C.
EU residency plus an existing Azure footprint point to Foundry EU, and an irreversible regulated action requires a human gate. Auto-deny (B) removes the mandatory gate; direct API (D) is constraint-blind to residency; a single aggregate target (E) hides per-claim-type failure (aggregate-metric).
During a 529 spike an active Fable 5.1 thinking session automatically falls back to a much older model and multi-turn behaviour becomes corrupted. What is the correct fix?
Show answer
Answer: B.
Thinking blocks are readable only by the producing model or newer, so fallback must preserve feature parity and log degraded mode. An iteration cap (A) and text parsing (D) don't address parity (prose-parsing anti-pattern); disabling retries (C) harms availability.
A team must launch in three weeks, lacks ops capacity, and the agent runtime is a commodity capability. What is the BEST build-vs-buy call?
Show answer
Answer: B.
No ops capacity plus a tight deadline plus a commodity capability points to managed/buy. Building (A) contradicts the constraints (over-engineered); delaying (C) fails the deadline (constraint-blind); a non-Claude product (D) abandons the requirement.
An agentic loop occasionally never terminates. Which is the CORRECT primary control plus backstop?
Show answer
Answer: B.
Control flow keys off
stop_reason; a cap is a backstop only. Text scanning (A) is the prose-parsing anti-pattern; a cap alone (C) is the iteration-cap-as-primary-stop anti-pattern; temperature (D) does not govern termination.A stem describes independent parallel investigations that must each explore with isolated context and then be synthesised. Which pattern is justified?
Show answer
Answer: B.
Parallel, context-isolated subtasks with fan-in are the textbook multi-agent signal. A single call (A) can't parallelise isolated exploration; a sequential shared-context workflow (C) doesn't isolate; fine-tuning (D) is unrelated to orchestration.
A design routes 100% to Opus 5 at 4× budget with acceptable quality. Which change is MOST cost-effective while preserving quality on hard cases?
Show answer
Answer: B.
A cascade routes cheap-first and escalates on validation failure, preserving quality on hard cases while most traffic runs cheaply; caching amortises the prefix. Blanket Haiku (A) sacrifices accuracy (constraint-blind); self-report (C) is anti-pattern #4; cutting users (D) doesn't address unit economics.
A stakeholder offers a vague goal and no numbers, and two teams disagree on scope. What is the BEST FIRST action?
Show answer
Answer: B.
Without criteria or constraints the design is unanchored; discovery surfaces the binding constraint and aligns stakeholders. Building (A), defaulting to a model (C), and inventing a target (D) all skip the anchoring step (constraint-blind).
Which change best completes an otherwise sound reference architecture that currently has no path from production signals back into the system?
Show answer
Answer: B.
A Professional-level architecture must close the loop from production signals into improvement; that is the feedback stage. A load balancer (A) is infrastructure; input validation (C) is the input stage; a bigger window (D) is unrelated to the missing loop.
A 4,000-token stable prefix on Sonnet 5 (
$2/MTok input) is reused at a 90% cache-hit rate over 10,000 req/day. Approximately how much of input cost does caching save?Show answer
Answer: B.
Read ≈ 0.1× input turns 4,000 prefix tokens into ~400 on hits, dropping daily input cost from ~
$90to ~$27(about 70%). Caching does help large stable prefixes (A); 50% is the Batch discount not caching (C); reads are cheap, not free (D).A pipeline must return strictly typed JSON for downstream systems and runs on Fable 5.1, where forcing tool use returns 400. What is the BEST approach?
Show answer
Answer: B.
Structured outputs with a
strictschema conform by construction and don't require forced tool use, which Fable 5.1 rejects with 400. Forcing tools then retrying (A) can't fix a 400; prompt-only JSON with regex (C) and free-text reparse (D) are brittle prose-parsing.Caching is enabled but the hit rate is ~3%. Investigation shows a per-request
request_idis prepended to the system prompt. What is the FIRST fix?Show answer
Answer: B.
A per-request token in the prefix changes it every call, so nothing caches; moving it out restores a stable prefix. Caching isn't broken (A); document length (C) only affects the minimum size; a bigger window (D) is unrelated.
A long agentic session on a thinking model overflows the window because tool results are huge. Which technique is correct, and what must be avoided?
Show answer
Answer: B.
Context editing removes stale tool results server-side and compaction summarises narrative, both avoiding thinking-block invalidation. Deleting turns client-side (A) breaks the append-only invariant; temperature (C) doesn't affect context size; in-turn self-summary (D) doesn't reclaim the window.
A team runs every request on Opus 5 at
effort: xhighwith acceptable quality and 5× budget; cost, not latency, is the complaint. Which TWO changes best cut cost while preserving quality?Show answer
Answer: A and B.
A cheap-first cascade and right-sized effort (re-validated per segment) cut cost while preserving quality on hard cases. Self-report (C) is anti-pattern #4; Fable 5.1 (D) is the most expensive model; removing evals (E) removes the quality guard (silent failure).
A hard rule 'never grant a discount above 20%' must be guaranteed in a Claude sales agent. Where does it belong?
Show answer
Answer: B.
Hard business rules require deterministic enforcement the model cannot be talked past. Prompt wording (A) and few-shot (C) are prompt-as-enforcement (anti-pattern #3); the thinking budget (D) is unrelated to enforcement.
After upgrading from Opus 4.6 to Opus 5, requests setting
budget_tokensfor thinking now return 400. What is the fix?Show answer
Answer: B.
budget_tokensis removed on Opus 5 (and Sonnet 5 / Fable 5.x) and returns 400; adaptive thinking pluseffortis the supported replacement. Retries (A) can't fix a 400; Haiku (C) is the only model still usingbudget_tokensbut is not an Opus equivalent; removing thinking (D) discards needed reasoning.A durable fact (a customer's contract tier) must persist across separate sessions without re-sending it in every prompt. Which mechanism fits, and what constraint applies?
Show answer
Answer: B.
The memory tool persists durable facts across sessions; the constraint is not to store secrets/PII inappropriately. Pasting per prompt (A) bloats context;
CLAUDE.local.md(C) is a Claude Code dev file, not a runtime store; relying on provider retention (D) is not a memory mechanism and raises compliance risk.A B2B assistant serves 300 tenants from one shared vector index; a tenant occasionally sees another tenant's document in an answer. What is the correct control?
Show answer
Answer: B.
Isolation is a retrieval-time security boundary: filter by
tenant_id/ACL before the vector search. A prompt rule (A) is bypassable prompt-as-enforcement; post-generation filtering (C) is too late because the chunk already entered context; a bigger model (D) doesn't isolate data.Across 5 queries the relevant chunk ranked 1, 4, 2, not-retrieved, and 1. What are recall@3 and MRR?
Show answer
Answer: B.
Three of five relevant chunks are in the top 3 (ranks 1, 2, 1) so recall@3 = 3/5 = 0.60; reciprocal ranks 1, 0.25, 0.5, 0, 1 give MRR = 2.75/5 = 0.55. Option A ignores the misses; C swaps the values; D is arithmetically wrong.
recall@10 = 0.63 (low) and answers frequently omit the needed fact. Where should the architect work FIRST?
Show answer
Answer: B.
Low recall means the right chunk isn't reaching the top-k, so the failure is upstream in retrieval. Grounding and citations (A, C) and a bigger generation model (D) can't help if the answer was never retrieved.
recall@10 = 0.95 but answers include facts absent from the retrieved chunks. Which layer failed and what is the fix?
Show answer
Answer: B.
High recall means the right context is present, so an unsupported fact is a grounding failure at generation. The retrieval-side fixes (A, C, D) target a healthy layer (recall-only reasoning).
A RAG system indexes with
text-embedding-3-largebut a new service queries with a different embedding model; similarity scores look random. What is the cause and fix?Show answer
Answer: B.
Embeddings from different models occupy different vector spaces, so cross-model similarity is meaningless; query and index must use the same embedding model. It isn't hardware (A); raising k (C) can't fix incompatible vectors; the generation model (D) is unrelated to retrieval similarity.
A knowledge assistant keeps giving the OLD answer, confidently, after a document was updated overnight. What should the architect investigate FIRST?
Show answer
Answer: B.
Confident-wrong immediately after a refresh points to stale retrieval/indexing; inspect the retrieved chunk IDs first. Prompt changes (A, D) and a bigger model (C) cannot fix stale retrieval.
A support agent has 22 tools, including
delete_ticketandrefundits role should never use. What is the correct remediation?Show answer
Answer: B.
Least privilege removes the capability so the excessive agency is gone. Logging (A) and confirmation (C) leave it reachable; a prompt rule (D) is prompt-as-enforcement and bypassable.
Contracts chunked at fixed 500 tokens produce answers citing the wrong sub-clause and losing surrounding context. Which TWO changes help MOST?
Show answer
Answer: A and B.
Structured documents need boundary-aware chunking and parent–child context. Dense-only (C), temperature (D), and removing citations (E) don't address chunking and the last two make it worse.
A remote MCP server over Streamable HTTP exposes powerful tools with no authentication. What MUST be added?
Show answer
Answer: B.
Remote MCP requires OAuth 2.1 and per-user authorization so the agent acts with the caller's authority. MCP isn't authenticated by default (A); prompts (C) and rate limits (D) don't address authorization.
After enabling backoff retries, some customers are charged twice. What is the correct fix?
Show answer
Answer: B.
Idempotency keys make retries on non-idempotent actions safe. Disabling retries (A) harms resilience; temperature (C) is irrelevant; weekly reconciliation (D) still double-charged customers (silent failure).
A 3M-document corpus changes daily and answers must cite the exact clause. Which approach is BEST?
Show answer
Answer: B.
Large, daily-changing, citation-requiring corpora are canonical RAG. Nightly fine-tuning (A) is an infeasible cadence for facts; the corpus exceeds/overspends context and loses citations (C); (D) inherits both problems.
An agent retrieves k=8 with dense-only and no reranker; precision is poor and exact part numbers are missed. Which TWO changes give the biggest quality lift?
Show answer
Answer: A and B.
Hybrid retrieval fixes exact-identifier misses and retrieve-wide-then-rerank fixes precision. Temperature (C) is irrelevant to retrieval; removing citations (D) harms grounding traceability; stuffing context (E) inflates cost without improving ranking (over-engineered).
On 2,000 paired items prompt B wins 1,080 (54%); on a separate 12-item hand test B won 7. Which result should drive promotion and why?
Show answer
Answer: B.
At n=2,000 the 80-win excess over 1,000 expected is ≈3.6σ (std dev ≈ 22.4), which is significant; 7/12 is coin-flip noise. The small test (A) is noise; A/B is reliable at scale (C); averaging win rates (D) is meaningless.
A prompt change wins significantly overall but Spanish-tier accuracy drops 6 points. What is the correct decision?
Show answer
Answer: B.
Per-segment no-regression is a hard gate; an aggregate win hiding a segment regression is anti-pattern #10. Promoting anyway (A, C) ships a known regression; dropping Spanish (D) re-hides it (aggregate-metric).
A team can only report a single overall accuracy number and cannot tell which segment fails. Which observability gap is the root cause?
Show answer
Answer: B.
Without segment tags and version fields only an aggregate is computable and regressions can't be attributed. GPU metrics (A) are irrelevant; judge size (C) doesn't create the reporting gap; latency (D) is a different signal.
Which TWO fields are MOST essential in a per-request eval record to support per-segment evaluation and A/B attribution?
Show answer
Answer: A and B.
Segment tags enable stratified metrics and prompt/model version enables attributing regressions and A/B results to a specific change. CPU temperature (C), an unlinked UUID (D), and a campaign name (E) don't support evaluation.
An eval grades answers with the same Sonnet 5 session that produced them and reports high scores. Why is this unsound and what is the FIRST fix?
Show answer
Answer: B.
Grading in the producing session retains the original bias (anti-pattern #9); independence plus human calibration is required. Cost (A) isn't the core issue; the judge needn't be larger (C); LLM-as-judge is valid when independent (D).
An interactive streaming UI feels slow even though total latency is acceptable. Which metric should be added?
Show answer
Answer: B.
In streaming UIs users perceive responsiveness by when tokens start (TTFT), not just total time. Total tokens (A), request count (C), and version count (D) are not latency-experience metrics.
A system fails only on the hardest 5% of cases and is fine elsewhere. What is the most likely cause and fix?
Show answer
Answer: B.
A clean 'only the hardest cases fail' signature points to model capability; a cascade escalates just those cases while keeping the cheap model elsewhere. Re-chunking (A) and rewriting the prompt (C) target working layers; replicas (D) don't affect correctness.
A team wants to cut cost 50% on a nightly bulk-summarisation job with no latency requirement. Which lever is BEST and what MUST follow?
Show answer
Answer: B.
A latency-tolerant bulk job fits Batch, and every cost change is followed by per-segment re-validation. Blind effort cuts (A) risk quality; changing interactive traffic (C) is out of scope; disabling evals (D) removes the safety net (silent failure).
Which TWO are the right metrics for a high-volume, cost-sensitive service with an interactive SLA?
Show answer
Answer: A and B.
Cost-sensitive plus interactive points to cost per task and p95 tail latency. Mean latency (C) hides the tail; total tokens (D) and version count (E) are not user-facing SLA metrics (aggregate-metric).
An LLM judge consistently rates the longer of two answers higher regardless of correctness. What is happening and the FIRST fix?
Show answer
Answer: B.
The judge exhibits length bias, fixed by human calibration and a correctness-focused rubric. The judge is not reliable (A); rewarding length (C) is the bug; deleting evals (D) abandons measurement.
A healthcare intake assistant handles PHI, is EU-based on AWS, and a vendor proposes Fable 5.1 with indefinite retention. Which TWO corrections are REQUIRED first?
Show answer
Answer: A and B.
A PHI/ZDR posture disqualifies Fable 5.1 and requires EU residency, a BAA, retention limits and erasure. 30-day deletion (C) is exactly what ZDR forbids; a prompt rule (D) can't protect PHI (prompt-as-enforcement); indefinite retention (E) breaches GDPR minimisation.
A US federal agency mandates FedRAMP High and wants the newest model the day it ships. What is the correct guidance?
Show answer
Answer: B.
FedRAMP High is provided through the Bedrock/Vertex boundary and the compliance requirement is binding over novelty. Direct API (A) is outside the boundary (constraint-blind); authorisation isn't inherent (C); a laptop (D) is unacceptable for federal data.
An incident-response runbook for an EU personal-data system is missing one legally required element. Which is it?
Show answer
Answer: B.
GDPR requires breach notification (generally within 72 hours), so the runbook must include it. Marketing (A), a dashboard list (C), and a cost projection (D) are not the legal requirement.
An agent summarising web pages reads hidden text instructing it to email internal data externally. What is it and the correct defence?
Show answer
Answer: A and B.
Malicious instructions in fetched content are indirect prompt injection; defence is untrusted-content handling plus least privilege and action validation. Trusting tool output (C) is the vulnerability; effort (D) is irrelevant; adding it to the prompt (E) executes the attack.
A support system escalates to a human whenever the customer 'sounds angry'. Why is this flawed?
Show answer
Answer: B.
Sentiment-based escalation (anti-pattern #5) conflates emotion with difficulty; a calm customer can have a complex, high-risk case. It is not optimal (A); frequency (C) isn't the core flaw; anger doesn't imply model failure (D).
A stem states 'high-risk processing of EU residents' personal data'. Which control belongs at the DESIGN gate?
Show answer
Answer: B.
GDPR high-risk processing requires a DPIA and residency/erasure decisions at the design gate. Waiting (A) is too late; a disclaimer (C) doesn't satisfy the DPIA; over-retention (D) breaches minimisation.
Which set best represents layered guardrails (defence in depth)?
Show answer
Answer: B.
Defence in depth stacks independent layers so a bypass of one is caught by the next. A single prompt (A), review-only (C), or regex-only (D) are single points of failure.
A prompt-injection incident is detected in production. What is a sound FIRST containment step?
Show answer
Answer: B.
Containment is fast rollback and disabling the exploited capability, scoped via traces. Deleting logs (A) destroys evidence; ignoring (C) prolongs harm; disclosing data (D) is a second breach.
Which TWO signal-word-to-control mappings are correct?
Show answer
Answer: A and B.
FedRAMP maps to the cloud boundary and a ZDR mandate disqualifies Fable 5.1. PHI needs a BAA and controls, not a prompt (C); EU erasure requires a deletion path, not indefinite retention (D); high-risk AI (EU AI Act) requires documentation, not skipping it (E).
Where should API keys and DB credentials for a Claude agent live?
Show answer
Answer: B.
Secrets belong in env/secret managers and never in model-visible config or logs. Putting them in the prompt (A), CLAUDE.md (C), or logs (D) creates exfiltration paths.
A VP sponsor was shown a 40-page technical ADR and remains unconvinced. What is the BEST corrective communication?
Show answer
Answer: B.
Sponsors need a decision matrix/cost model/one-page brief at their altitude, not an engineering ADR. Resending the ADR (A) repeats the wrong-altitude error; raw eval logs (C) and the runbook (D) are the wrong artefacts for a sponsor.
Using a weighted decision matrix (SLA 0.25, cost 0.25, reliability 0.20, flexibility 0.15, effort 0.15), Workflow scores 5,5,5,3,4 and Agentic scores 3,3,3,5,3. Which wins and why?
Show answer
Answer: B.
Workflow = 0.25·5+0.25·5+0.20·5+0.15·3+0.15·4 = 4.60; Agentic = 0.25·3+0.25·3+0.20·3+0.15·5+0.15·3 = 3.30, so workflow wins on the business-weighted criteria. Flexibility (A) is only 0.15; it isn't a tie (C); the matrix decides transparently (D).
A sponsor asks for 'faster support'. How should this become an SLA?
Show answer
Answer: B.
SLAs must be measurable and per-segment, mirroring eval criteria, with regular reporting. Vague promises (A) can't be verified; mean-only (C) hides the tail; undefined (D) invites disputes.
Ops refuses a handoff. Which deliverable set satisfies the handoff exit criterion?
Show answer
Answer: B.
Handoff's exit criterion is operability and recoverability, provided by the runbook plus architecture/data-flow docs. A deck (A), code alone (C), or a cost model (D) don't let operators run and recover it.
A model a service pins was retired. Which TWO steps make the migration correct?
Show answer
Answer: A and B.
Deprecation is a governed lifecycle event: assess breaking changes, re-validate per segment, canary with rollback, and communicate. A blind swap (C) risks regressions; calling a retired model (D) fails; disabling evals (E) removes the safety net.
During discovery a stakeholder says 'make onboarding faster with AI' and offers no numbers. What is the FIRST step?
Show answer
Answer: B.
Discovery converts vague asks into numeric, per-segment criteria with constraints and sign-off before design. Building (A), assuming a target (C), or escalating (D) skip the anchoring step (constraint-blind).
What is the correct distinction among SLA, SLO and SLI?
Show answer
Answer: B.
SLI (indicator) → SLO (internal objective) → SLA (external commitment). They are not synonyms (A), the internal/external roles are not swapped (C), and the SLI is a metric, not a contract (D).
A technically strong system sees low adoption. Which TWO change-management levers help MOST?
Show answer
Answer: A and B.
Adoption is driven by enablement and phased rollout with feedback and demonstrated value. Forced use (C), removing gates (D), and hiding limits (E) backfire.
Which artefact best communicates residual risk (after mitigation) to a compliance stakeholder?
Show answer
Answer: B.
A risk register with likelihood/impact, owners, mitigations and residual risk is the compliance-facing artefact. An ADR (A) records a technical decision; a cost model (C) is economics; a code diagram (D) is for engineers.
A platform team must guarantee
git push --forceis blocked for every developer and unbypassable. Where does this belong?Show answer
Answer: B.
Deterministic, non-bypassable blocking is a
PreToolUsehook (exit 2) enforced through managed policy. A CLAUDE.md sentence (A) is prose guidance (prompt-as-enforcement);CLAUDE.local.md(C) is personal/overridable; Slack (D) is not enforcement.A team wants Claude to review diffs automatically in CI. What is the correct setup?
Show answer
Answer: B.
CI is non-interactive, so headless mode with structured output and restricted tools is correct. Interactive use (A) can't run in CI; all-tools (C) and disabled permissions (D) violate least privilege.
A platform group is rolling Claude Code to 15 teams and wants low risk. Which TWO choices are BEST?
Show answer
Answer: A and B.
Phased rollout plus managed-policy enforcement and a shared checked-in catalogue is the safe pattern. A day-one mandate (C) skips trust-building; unvetted MCP servers (D) break governance; lines of code (E) is a vanity metric.
Last updated Sep 18, 2026