CCDV-F Practice Exam 2
A second full-length, blueprint-weighted practice exam with new, harder items.
This is a second full-length CCDV-F practice exam with 53 entirely new items — none repeat Practice Exam 1 or the domain-page questions. It is pitched slightly harder than Exam 1: more multi-constraint scenarios, more FIRST / MOST cost-effective / TWO qualifiers, and distractors that are individually plausible until you apply the binding constraint. The distribution matches the blueprint exactly, so your per-domain score here is a realistic booking-readiness signal.
Instructions
- 53 items · 120 minutes. Budget roughly 2 minutes per item; flag and return to the hard ones.
- Scoring is scaled 100–1000, pass 720. As a study heuristic, aim for ≥ 80% raw (≈ 43/53) before booking.
- Items are multiple-choice (select one) or multiple-response (select two — the item says so). There is no guessing penalty, so answer everything.
- Watch the qualifiers:
FIRSTasks for the initial correct action,MOST cost-effectiveranks on cost given the constraints, andTWOmeans exactly two correct options.
Domain distribution
| Domain | Weight | Items here |
|---|---|---|
| D1 · Applications and Integration | 33.1% | 18 |
| D2 · Model Selection and Optimization | 16.8% | 9 |
| D3 · Agents and Workflows | 14.7% | 9 |
| D4 · Prompt and Context Engineering | 11.0% | 6 |
| D5 · Tools and MCPs | 10.6% | 5 |
| D6 · Security and Safety | 8.1% | 4 |
| D7 · Claude Code | 3.1% | 1 |
| D8 · Eval, Testing and Debugging | 2.6% | 1 |
How to use both exams
- Sit Practice Exam 1 first. Use it as a diagnostic to find weak domains, then restudy those domain pages before returning.
- Use Practice Exam 2 as the booking gate. Its harder, multi-constraint stems are closer to the pressure of the real exam. Score ≥ 80% raw here, with no single heavy domain (D1/D2/D3) badly lagging, before you book.
- Compare per-domain results across both exams. A domain that is strong on Exam 1 but weak on Exam 2 usually means you learned the recall facts but not the multi-constraint reasoning — go back to that domain’s scenario walkthrough and misconceptions table.
Score interpretation
| Raw score (of 53) | Reading |
|---|---|
| ≥ 47 (≈ 89%) | Strong — comfortably above the likely pass bar even on harder items |
| 43–46 (≈ 81–87%) | On track — book once no heavy domain lags |
| 38–42 (≈ 72–79%) | Borderline — revisit weak domains, especially D1/D2/D3 |
| < 38 (< 72%) | Not ready — restudy the heavy domains and re-test |
Take the practice exam
The interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
53 questions · one at a time · 120-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A production integration streams responses and must (a) render text as it arrives and (b) record the final
stop_reasonand token usage for billing. Which SSE handling is correct?Show answer
Answer: A.
Text streams as
text_deltainsidecontent_block_delta; the finalstop_reasonand usage arrive onmessage_delta(beforemessage_stop). B is recall-only and wrong:message_starthas no final usage. C misplacesstop_reasononmessage_stop. D is silent-failure thinking:message_stopmerely ends the stream and carries no text.A tool-use loop appends the assistant's
tool_useblocks and thetool_resultblocks, but intermittently returns400 invalid_request. Logs show occasional turns where atool_resultis missing for one of two paralleltool_useblocks. What is the FIRST fix?Show answer
Answer: A.
Every
tool_useblock needs a matchingtool_result(same id) in the next user turn; a missing one causes the 400. A fixes the root cause. B is wrong because a 400 is deterministic, not transient. C (temperature) does not guarantee serial calls and is prompt-as-enforcement thinking. D over-engineers by removing a valid capability instead of returning all results.A service reuses a 30,000-token system+schema prefix on Sonnet 5 across ~500 calls/hour, each with a short unique question. Costs are dominated by input tokens. Which TWO changes cut input cost the MOST while keeping output quality unchanged?
Show answer
Answer: A and B.
Prompt caching a stable prefix and reusing it within the TTL drops the prefix cost to ~10% on hits. A and B are the levers. C (temperature) and D (
max_tokens) do not affect input caching cost. E is wrong: the cacheable prefix must come first; putting it after the variable question prevents cache hits.A worker calls the SDK in a tight
asyncio.gatherover 5,000 items and starts seeing many429responses with aretry-afterheader. Which change is MOST cost-effective and correct for this latency-tolerant workload?Show answer
Answer: A.
A latency-tolerant bulk job belongs on the Batches API: half the cost and no self-inflicted rate-limit storm. A is correct. B ignores
retry-afterand worsens the limit. C is constraint-blind: account limits are not multiplied by keys and may violate terms. D over-engineers with a costlier model that does not address rate limits.With adaptive thinking enabled on Opus 5, code reads
response.content[0].textand throwsAttributeErroron some responses but not others. What is the root cause and correct handling?Show answer
Answer: A.
When thinking or tools are active, the first content block may be
thinkingortool_use, so index 0 is not always text. A iterates and filters by type. B is wrong: thinking does not suppress text and disabling it is over-engineering. C misstates the shape (content is a typed array, not a string). D is unrelated to response shape.An enterprise must keep all inference inside its Google Cloud project under existing IAM and VPC controls. Which client and configuration is appropriate?
Show answer
Answer: A.
Vertex keeps inference in the customer's GCP project under IAM. A is correct. B is AWS, the wrong cloud. C is constraint-blind: storing the key in Secret Manager still calls the external Anthropic API, leaving GCP. D is a consumer surface, not an in-account programmatic path.
A batch of 20,000 requests is created and the code immediately calls
results(), which returns nothing, so it recreates the batch and doubles the spend. What is the correct lifecycle?Show answer
Answer: A.
Batches are asynchronous: poll to
ended, then read results keyed bycustom_id, and use idempotency on create. A is correct. B is wrong (batches are not synchronous; recreating duplicates work). C hammers the API pointlessly. D invents a nonexistent model constraint.A response returns
stop_reason: 'refusal'. The current code catches it in a genericexcept, logs 'API error', and retries with backoff. What is wrong and what should happen?Show answer
Answer: A.
refusalis a deliberate safety stop reason and should be handled by policy, not retried. A is correct and also fixes the generic-error anti-pattern. B blind-retries a non-transient signal. C confuses it withmax_tokens. D confuses it withtool_use.A team hardens a new integration before production. Which TWO practices belong to the hardening stage rather than prompt wording?
Show answer
Answer: A and B.
Hardening is about reliability: backoff on transient errors and schema validation-retry. A and B are correct. C and D are cosmetic prompt tweaks, not hardening. E is self-report/same-session review, an anti-pattern that does not harden the integration.
A 180-page contract PDF is reused across dozens of extraction requests per day. Base64 bytes are re-sent every call, dominating bandwidth and latency. What is the MOST cost-effective input method?
Show answer
Answer: A.
The Files API uploads once and references by
file_id, eliminating repeated uploads. A is correct. B still re-sends large text each call. C is even heavier (images per page every time). D (streaming) affects delivery, not the repeated upload cost.A caching layer marks a 12,000-token prefix with
cache_control, but cache hit rate is near zero. Inspection shows the system prompt embeds the current UTC timestamp. Which statement is correct?Show answer
Answer: A.
A cache hit requires a byte-identical prefix; a per-request timestamp changes it every call. A fixes it by keeping variable content after the breakpoint. B invents a size cap. C misdiagnoses (12k is above the ~1024 minimum). D is wrong: a longer TTL still requires identical bytes to hit.
Under sustained load the app hits
429on the ITPM axis first, even though RPM is fine. Which change addresses the actual limiting axis MOST directly?Show answer
Answer: A.
The limit hit is input-tokens-per-minute, so cutting input tokens (caching, trimming) or raising the tier is the direct fix. A is correct. B ignores the binding axis. C raises OUTPUT tokens, worsening a different axis. D (temperature) does not change token counts.
A developer sends
messagesbeginning with arole: 'assistant'turn and gets400 invalid_request. What is the rule?Show answer
Answer: A.
Conversations start with
userand alternate. A states the rule. B is wrong (a 400 is not a 429). C is wrong:systemis a top-level field, not a message entry on current models. D invents a streaming constraint.A support engineer asks for the single most useful thing to include in application logs so Anthropic support can trace a specific failed call. What is it?
Show answer
Answer: A.
The
request-idis the correlation handle support and your traces both key on. A is correct. B logs a secret (security violation). C logs PII (hygiene violation). D is insufficient and not a trace handle.A harness on Fable 5.1 trims history by editing and reordering earlier turns; users report errors about invalid thinking blocks. Which redesign is correct?
Show answer
Answer: A.
Fable 5.1 invalidates later thinking when earlier turns change, so the harness must be append-only and trim server-side. A is correct. B blind-retries a deterministic design error. C is impossible (Fable 5.1 always thinks). D does not stop thinking blocks from being produced.
A migration keeps the same Messages API code but must run through Amazon Bedrock for compliance. Which TWO things actually change versus the direct Anthropic API?
Show answer
Answer: A and B.
Across Bedrock/Vertex the request shape and
stop_reasonsemantics are the same; what changes is the client/auth and the model ID format. A and B are correct. C, D and E are unchanged — the whole point is portability of the request body and control fields.A code path calls
client.messages.createwithmax_tokensomitted, expecting a sensible default, and receives400. Which is the correct explanation?Show answer
Answer: A.
max_tokensis a required field capping generated output. A is correct. B misreads a deterministic 400 as transient. C confuses it with the context window. D invents a tools-only condition.An SDK client is created once and shared across an async web server handling thousands of concurrent requests. Which configuration best balances reliability and correctness?
Show answer
Answer: A.
An async client with a timeout, bounded retries, and a concurrency cap is the robust pattern. A is correct. B blocks the loop with a sync client and no timeout. C removes safety nets for transient errors. D ignores rate limits and causes 429 storms.
A pipeline runs 200,000 short sentiment classifications overnight; latency is irrelevant and the label set is simple. Which choice is MOST cost-effective?
Show answer
Answer: A.
Simple, high-volume, latency-tolerant work maps to the cheapest capable model plus batching. A is correct. B and D are over-engineered and costly for trivial labels. C (streaming) does not reduce token cost and is irrelevant to an offline job.
A team wants a fixed reasoning budget of 1,500 thinking tokens. On which current model can they set
budget_tokens, and what should the others use?Show answer
Answer: A.
budget_tokensis Haiku-4.5-only; other current models use adaptive thinking with effort. A is correct. B is false (it returns 400 elsewhere). C is false (the others do support thinking, just adaptive). D invents a pairing requirement.A task sends 4,000 input + 800 output tokens, 30,000 times/day. Sonnet 5 ($2/$10 per MTok) and Haiku 4.5 ($1/$5) both clear the quality bar. What is the daily cost difference, and which is cheaper?
Show answer
Answer: A.
Haiku: input 4000*30000*$1/1e6 = $120 + output 800*30000*$5/1e6 = $120 = $240/day. Sonnet: $240 input + $240 output = $480/day. A is correct (~$240 difference). B ignores that per-token prices differ. C invents a batching effect. D reverses the arithmetic.
Most traffic is trivial; a small tail needs deep reasoning; the hard cases must not degrade and cost must stay low. A junior proposes escalating whenever Claude self-reports confidence below 0.8. What is the correct design?
Show answer
Answer: A.
A cascade with a programmatic escalation gate keeps the bulk cheap without hurting the hard tail. A is correct. B is over-engineered/costly; C fails the hard cases (constraint-blind); D relies on self-reported confidence (anti-pattern #4).
Users say a chat UI 'feels sluggish' although total task time is acceptable, and output quality must not change. Which TWO levers best improve the perceived experience?
Show answer
Answer: A and B.
Perceived latency is time-to-first-token; streaming plus prefix caching (skipping prefix reprocessing) both get output flowing sooner without changing quality. A and B are correct. C, D and E all add latency and/or cost and change behaviour.
For a fair offline comparison of two prompt versions, a developer runs each once at
temperature: 0on a floating model alias and treats identical scores as proof of equivalence. What is the flaw?Show answer
Answer: A.
Temperature 0 lowers but does not eliminate variance, and a floating alias undermines reproducibility. A is correct. B is the misconception being tested. C (effort) adds variability, not rigour. D (streaming) is a delivery mechanism unrelated to reproducibility.
Which TWO statements about context windows and output caps across the current lineup are correct?
Show answer
Answer: A and B.
Sonnet 5 and Opus 5 are 1M context; Haiku 4.5 is 200k. A and B are correct. C is wrong (Haiku is 200k). D confuses Fable 5.1's 128k MAX OUTPUT with its 1M context. E understates Sonnet 5's 1M window.
A 10,000-document extraction runs overnight and is cost-sensitive; Sonnet 5 clears the quality bar. Each request has a 2,000-token stable prefix and 1,000 unique tokens plus 500 output tokens. Which combination is MOST cost-effective?
Show answer
Answer: A.
A latency-tolerant job should stack all three levers: cheapest model that clears the bar, Batches (50% off), and caching the stable prefix. A is correct. B is over-engineered. C leaves the 50% batch discount unused. D is invalid because Haiku 4.5 has no
effortparameter.A team migrating from Sonnet 5 to Opus 5 wants to minimise risk. Which sequence is correct?
Show answer
Answer: A.
Centralised, flagged, eval-gated, gradually ramped migration is the safe path. A is correct. B risks a wide blast radius; C causes model drift and inconsistency; D changes production with no safety net.
A workflow has three fixed, ordered stages (transcribe → summarise → classify), each with known inputs and outputs. A developer proposes an autonomous agent that decides the steps. What is the BEST design and why?
Show answer
Answer: A.
Fixed ordered steps are the definition of prompt chaining. A is correct. B is over-engineered — autonomy adds unpredictability with no benefit here. C sacrifices control and testability. D is for tasks whose subtasks are unknown at design time, not this one.
An agent loop occasionally never terminates. The current code stops when the assistant text contains 'done' and also caps at 5 iterations. What is the correct PRIMARY termination?
Show answer
Answer: A.
Termination must be driven by
stop_reason; the cap is a backstop. A is correct. B makes the cap primary (anti-pattern #2). C parses prose (anti-pattern #1). D does not create a reliable control signal.A research agent's context fills with verbose intermediate tool output, degrading later reasoning and inflating cost. Which design fixes this MOST directly?
Show answer
Answer: A.
Isolated-context subagents keep noisy intermediate work out of the coordinator's window and return only summaries. A is correct. B caps output length, not context bloat. C does not shrink tool output. D removes needed capability (constraint-blind).
In the Claude Agent SDK, which TWO options directly enforce least privilege and deterministic safety for an agent that must never force-push or delete data?
Show answer
Answer: A and B.
An allowlist plus a blocking hook enforce least privilege and hard stops deterministically. A and B are correct. C weakens the backstop. D is prompt-as-enforcement (anti-pattern #3). E removes safeguards entirely.
A team wants Anthropic to host the agent loop and sandbox so they can ship standard agentic tasks fast with minimal ops. Which option fits, and what is the main trade-off?
Show answer
Answer: A.
Managed Agents host the loop and sandbox with lower environment control. A is correct. B misstates the SDK (you host it). C provides no managed loop/sandbox and you own termination. D is not an agent at all.
A single agent is configured with 18 tools and frequently selects the wrong one. Which remedy aligns with Anthropic guidance?
Show answer
Answer: A.
Too many tools is anti-pattern #8; the fix is fewer active tools, subagent specialisation, and tool search +
defer_loading. A is correct. B worsens overload. C is brittle and breaks on Fable 5.1. D does not improve selection.An agent loop returns
stop_reason == 'max_tokens'mid-answer, and the harness marks the turn complete. What is the correct handling?Show answer
Answer: A.
max_tokensmeans truncation, not completion. A recovers the missing content. B silently truncates work (silent-failure). C regenerates rather than continuing and loses the partial. D changes the model without recovering truncated content.A task must run several independent lookups and then aggregate them; the results do not depend on each other. Which pattern fits BEST?
Show answer
Answer: A.
Independent concurrent subtasks with aggregation is parallelization. A is correct. B serialises independent work needlessly. C is for iterative quality improvement. D dispatches one category, not many concurrent subtasks.
A research coordinator delegates topics to worker subagents. Which TWO choices keep the coordinator's context clean and costs down?
Show answer
Answer: A and B.
Isolated worker contexts plus distilled-summary returns keep the coordinator window clean and cheap. A and B are correct. C and D reintroduce bloat; E is anti-pattern #8 (too many tools).
An endpoint on Sonnet 5 must return schema-conforming JSON on every call. Which approach is MOST reliable?
Show answer
Answer: A.
Schema-constrained structured output plus validation-retry is the reliable pattern. A is correct. B is prompt-only and drifts. C addresses truncation, not conformance. D omits validation (silent-failure risk).
A validation-retry loop rejects the model's JSON but re-sends only the original prompt, so the same error recurs until attempts run out. What is the fix?
Show answer
Answer: A.
Effective validation-retry feeds back the concrete error so the model can correct it. A is correct. B blindly repeats the same mistake (recall-only). C changes the model without telling it what was wrong. D can worsen truncation and conveys nothing.
A long-running agent slows and loses focus as history grows. The narrative must stay coherent, but old tool results are large and stale. Which TWO server-side controls address this correctly?
Show answer
Answer: A and B.
Context editing clears stale tool outputs; compaction condenses the narrative while keeping the thread. A and B are correct. C caps output length only. D and E do nothing for bloat/drift.
A long-document QA prompt places the user's question first, then 200 pages of source, and answers are poorly grounded while caching never helps. Which single reordering fixes both problems?
Show answer
Answer: A.
Documents-first, query-last improves grounding and makes the stable prefix cacheable in one change. A is correct. B still precedes the evidence with the query. C and D harm coherence and caching.
A team needs reliable structured extraction on Fable 5.1, but forcing a specific tool with
tool_choicereturns400. What should they use instead?Show answer
Answer: A.
Fable 5.1 rejects
any/forced tool choice, so structured outputs orstrict/autoare the paths. A is correct. B retries a deterministic 400. C is impossible and irrelevant. D does not influence tool choice.A summariser handles common documents well but repeatedly mishandles a rare handwritten-form format. Which TWO prompt changes best improve it?
Show answer
Answer: A and B.
Edge-case failures call for representative edge-case examples and greater diversity. A and B are correct. C adds more of the already-handled case (recall-only). D adds variability. E discards the steering entirely.
Claude frequently calls the wrong tool and passes malformed arguments for a closed-set field like priority. What is the FIRST and MOST effective fix?
Show answer
Answer: A.
The description and schema are the primary levers for correct tool use; an
enumprevents invented values. A is correct. B, C and D are secondary and do not address a vague description or loose schema.A
get_ordertool times out during a parallel tool call alongside a successfulsearch_kb. Which TWO actions keep the protocol valid and let Claude recover?Show answer
Answer: A and B.
Every
tool_useneeds a matchingtool_resultin the same turn; errors useis_error: true. A and B are correct. C causes a 400 and hides the failure (silent-failure). D breaks id matching. E abandons the run and loses diagnostics.Which TWO capabilities are server-side (Anthropic-hosted) built-in tools rather than client-side custom tools?
Show answer
Answer: A and B.
Web search and code execution are Anthropic-hosted server-side tools. A and B are correct. C, D and E run in your code/infrastructure and are client-side custom tools.
A team wants the SAME tools available in Claude Code, Claude Desktop and their production Messages API service, maintained in one place, with a remote server secured properly. What is the best mechanism and its security requirement?
Show answer
Answer: A.
MCP is the cross-host reuse mechanism; remote servers use OAuth 2.1. A is correct. B duplicates maintenance. C is Claude-Code-only packaged know-how, not a cross-host tool server. D holds instructions, not executable tools.
In MCP, which primitive is model-controlled (Claude decides when to invoke it), and how do the other two differ?
Show answer
Answer: A.
Tools (model), resources (application), prompts (user) is the correct control mapping. A is correct. B and C scramble the mapping; transports (stdio/HTTP) are not a primitive. D is false — only tools are model-controlled.
An agent summarises fetched web pages. One page hides an HTML comment instructing Claude to call
transfer_funds, and the agent nearly complies. Which combination BEST prevents the transfer?Show answer
Answer: A.
This is indirect injection; the defence is content boundaries (treat external text as data) plus a deterministic hook on the dangerous tool. A is correct. B is prompt-as-enforcement (anti-pattern #3). C and D do not stop injection — temperature and model size are not injection defences.
A read-only support agent currently has
get_order,search_kb,issue_refundanddelete_accounttools. Which TWO changes best satisfy least privilege and secrets hygiene?Show answer
Answer: A and B.
Least privilege removes unneeded powerful tools; secrets live in env/secret manager. A and B are correct. C only detects harm after the fact; D leaks the key into version control; E is prompt-as-enforcement (anti-pattern #3).
A healthcare app must process PHI with zero data retention and FedRAMP High. Which TWO decisions fit the constraints?
Show answer
Answer: A and B.
Bedrock/Vertex provide FedRAMP High and BAA coverage; a ZDR-eligible model with PHI minimisation meets zero-retention. A and B are correct. C breaks ZDR (Fable 5.1 mandates 30-day retention). D over-shares PHI. E removes enforcement/auditability.
In a layered-guardrail design (detect → instruct → enforce → verify → approve), where must 'no production deletes without approval' actually be enforced?
Show answer
Answer: A.
Critical, irreversible rules must live in the enforce layer as deterministic hooks. A is correct. B is prompt-as-enforcement (anti-pattern #3). C detects but does not block. D checks after the action has already run.
A team must block the
rm -rfcommand family for everyone in Claude Code, in a way no developer can override, while still running non-interactively in CI. Which TWO steps achieve this?Show answer
Answer: A and B.
A managed
denypattern blocks the family non-overridably, and headless-pwith a tool allowlist runs safely in CI. A and B are correct. C is prose guidance, not enforcement. D removes safeguards. E fires after the damage is done.A model scores 93% overall on evals yet a production incident involves handwritten forms, and the eval pipeline had the generator grade its own answers in the same session. Which TWO practices would have caught and correctly judged this?
Show answer
Answer: A and B.
Per-segment gating surfaces the handwritten-forms failure (anti-pattern #10), and a separate-session/different-model judge avoids self-review bias (anti-pattern #9). A and B are correct. C is same-session self-review (#9); D is aggregate-masking (#10); E removes the safety net entirely.
Last updated Sep 18, 2026