AI Cert Prep
Type to search documentation.

CCDV-F Practice Exam 2

A second full-length, blueprint-weighted practice exam with new, harder items.

This is a second full-length CCDV-F practice exam with 53 entirely new items — none repeat Practice Exam 1 or the domain-page questions. It is pitched slightly harder than Exam 1: more multi-constraint scenarios, more FIRST / MOST cost-effective / TWO qualifiers, and distractors that are individually plausible until you apply the binding constraint. The distribution matches the blueprint exactly, so your per-domain score here is a realistic booking-readiness signal.

Instructions

  • 53 items · 120 minutes. Budget roughly 2 minutes per item; flag and return to the hard ones.
  • Scoring is scaled 100–1000, pass 720. As a study heuristic, aim for ≥ 80% raw (≈ 43/53) before booking.
  • Items are multiple-choice (select one) or multiple-response (select two — the item says so). There is no guessing penalty, so answer everything.
  • Watch the qualifiers: FIRST asks for the initial correct action, MOST cost-effective ranks on cost given the constraints, and TWO means exactly two correct options.

Domain distribution

DomainWeightItems here
D1 · Applications and Integration33.1%18
D2 · Model Selection and Optimization16.8%9
D3 · Agents and Workflows14.7%9
D4 · Prompt and Context Engineering11.0%6
D5 · Tools and MCPs10.6%5
D6 · Security and Safety8.1%4
D7 · Claude Code3.1%1
D8 · Eval, Testing and Debugging2.6%1

How to use both exams

  • Sit Practice Exam 1 first. Use it as a diagnostic to find weak domains, then restudy those domain pages before returning.
  • Use Practice Exam 2 as the booking gate. Its harder, multi-constraint stems are closer to the pressure of the real exam. Score ≥ 80% raw here, with no single heavy domain (D1/D2/D3) badly lagging, before you book.
  • Compare per-domain results across both exams. A domain that is strong on Exam 1 but weak on Exam 2 usually means you learned the recall facts but not the multi-constraint reasoning — go back to that domain’s scenario walkthrough and misconceptions table.

Score interpretation

Raw score (of 53)Reading
≥ 47 (≈ 89%)Strong — comfortably above the likely pass bar even on harder items
43–46 (≈ 81–87%)On track — book once no heavy domain lags
38–42 (≈ 72–79%)Borderline — revisit weak domains, especially D1/D2/D3
< 38 (< 72%)Not ready — restudy the heavy domains and re-test

Take the practice exam

The interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

53 questions · one at a time · 120-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · Applications and IntegrationSelect one

    A production integration streams responses and must (a) render text as it arrives and (b) record the final stop_reason and token usage for billing. Which SSE handling is correct?

    • A. Read text_delta from content_block_delta for display, and read stop_reason and final usage from the message_delta event.
    • B. Read everything from message_start, which contains the full text and final usage.
    • C. Accumulate content_block_start events; stop_reason is on message_stop.
    • D. Only message_stop carries text, usage and stop_reason.
    Show answer

    Answer: A.

    Text streams as text_delta inside content_block_delta; the final stop_reason and usage arrive on message_delta (before message_stop). B is recall-only and wrong: message_start has no final usage. C misplaces stop_reason on message_stop. D is silent-failure thinking: message_stop merely ends the stream and carries no text.

  2. Q2D1 · Applications and IntegrationSelect one

    A tool-use loop appends the assistant's tool_use blocks and the tool_result blocks, but intermittently returns 400 invalid_request. Logs show occasional turns where a tool_result is missing for one of two parallel tool_use blocks. What is the FIRST fix?

    • A. Return exactly one tool_result per tool_use block, all in the same following user turn, each keyed to its tool_use_id.
    • B. Retry the 400 with exponential backoff and jitter.
    • C. Lower temperature so Claude emits fewer parallel tool calls.
    • D. Set disable_parallel_tool_use: true permanently to avoid the problem.
    Show answer

    Answer: A.

    Every tool_use block needs a matching tool_result (same id) in the next user turn; a missing one causes the 400. A fixes the root cause. B is wrong because a 400 is deterministic, not transient. C (temperature) does not guarantee serial calls and is prompt-as-enforcement thinking. D over-engineers by removing a valid capability instead of returning all results.

  3. Q3D1 · Applications and IntegrationSelect two

    A service reuses a 30,000-token system+schema prefix on Sonnet 5 across ~500 calls/hour, each with a short unique question. Costs are dominated by input tokens. Which TWO changes cut input cost the MOST while keeping output quality unchanged?

    • A. Place the stable prefix first and mark its end with cache_control so it is written once and read at ~0.1x thereafter.
    • B. Reuse the identical cached prefix on every call within the TTL.
    • C. Set temperature: 0.
    • D. Raise max_tokens to reduce truncation retries.
    • E. Move the prefix after the user question so it caches better.
    Show answer

    Answer: A and B.

    Prompt caching a stable prefix and reusing it within the TTL drops the prefix cost to ~10% on hits. A and B are the levers. C (temperature) and D (max_tokens) do not affect input caching cost. E is wrong: the cacheable prefix must come first; putting it after the variable question prevents cache hits.

  4. Q4D1 · Applications and IntegrationSelect one

    A worker calls the SDK in a tight asyncio.gather over 5,000 items and starts seeing many 429 responses with a retry-after header. Which change is MOST cost-effective and correct for this latency-tolerant workload?

    • A. Move the 5,000 items to the Message Batches API (50% discount, results within 24h) instead of hand-rolled concurrency.
    • B. Ignore retry-after and retry immediately in a tighter loop.
    • C. Rotate through several API keys to multiply the account limit.
    • D. Switch every call to Opus 5 to reduce the number of retries.
    Show answer

    Answer: A.

    A latency-tolerant bulk job belongs on the Batches API: half the cost and no self-inflicted rate-limit storm. A is correct. B ignores retry-after and worsens the limit. C is constraint-blind: account limits are not multiplied by keys and may violate terms. D over-engineers with a costlier model that does not address rate limits.

  5. Q5D1 · Applications and IntegrationSelect one

    With adaptive thinking enabled on Opus 5, code reads response.content[0].text and throws AttributeError on some responses but not others. What is the root cause and correct handling?

    • A. content[0] is sometimes a thinking (or tool_use) block; iterate content and concatenate blocks where type == 'text'.
    • B. Thinking randomly disables text output; disable thinking to stabilise it.
    • C. The response content is sometimes a plain string; call .strip() first.
    • D. Text only appears when streaming; switch to streaming.
    Show answer

    Answer: A.

    When thinking or tools are active, the first content block may be thinking or tool_use, so index 0 is not always text. A iterates and filters by type. B is wrong: thinking does not suppress text and disabling it is over-engineering. C misstates the shape (content is a typed array, not a string). D is unrelated to response shape.

  6. Q6D1 · Applications and IntegrationSelect one

    An enterprise must keep all inference inside its Google Cloud project under existing IAM and VPC controls. Which client and configuration is appropriate?

    • A. AnthropicVertex on Google Vertex AI with Vertex model IDs and region scoping.
    • B. AnthropicBedrock with AWS SigV4.
    • C. The default Anthropic client with the API key stored in Google Secret Manager.
    • D. claude.ai with Google SSO enabled.
    Show answer

    Answer: A.

    Vertex keeps inference in the customer's GCP project under IAM. A is correct. B is AWS, the wrong cloud. C is constraint-blind: storing the key in Secret Manager still calls the external Anthropic API, leaving GCP. D is a consumer surface, not an in-account programmatic path.

  7. Q7D1 · Applications and IntegrationSelect one

    A batch of 20,000 requests is created and the code immediately calls results(), which returns nothing, so it recreates the batch and doubles the spend. What is the correct lifecycle?

    • A. Poll processing_status until it is ended, then stream results() and match each by custom_id; use an idempotency key on create to avoid duplicates.
    • B. Batches are synchronous; an empty result means the batch failed, so recreate it.
    • C. Call results() in a tight loop with no delay until data appears.
    • D. Batches only return results on Opus 5; switch models and retry.
    Show answer

    Answer: A.

    Batches are asynchronous: poll to ended, then read results keyed by custom_id, and use idempotency on create. A is correct. B is wrong (batches are not synchronous; recreating duplicates work). C hammers the API pointlessly. D invents a nonexistent model constraint.

  8. Q8D1 · Applications and IntegrationSelect one

    A response returns stop_reason: 'refusal'. The current code catches it in a generic except, logs 'API error', and retries with backoff. What is wrong and what should happen?

    • A. refusal is a safety decline, not an error; route it to a policy/human path and surface diagnostic context — do not blind-retry.
    • B. refusal is transient; the backoff retry is correct.
    • C. refusal means the output was truncated; raise max_tokens and retry.
    • D. refusal means a tool was requested; execute the tool.
    Show answer

    Answer: A.

    refusal is a deliberate safety stop reason and should be handled by policy, not retried. A is correct and also fixes the generic-error anti-pattern. B blind-retries a non-transient signal. C confuses it with max_tokens. D confuses it with tool_use.

  9. Q9D1 · Applications and IntegrationSelect two

    A team hardens a new integration before production. Which TWO practices belong to the hardening stage rather than prompt wording?

    • A. Retry 429/5xx/529 with exponential backoff + jitter, honouring retry-after.
    • B. Validate structured output against a schema and retry with the error on failure.
    • C. Make the answer friendlier in tone.
    • D. Add more emoji to the system prompt.
    • E. Ask the model to double-check its own work in the same call.
    Show answer

    Answer: A and B.

    Hardening is about reliability: backoff on transient errors and schema validation-retry. A and B are correct. C and D are cosmetic prompt tweaks, not hardening. E is self-report/same-session review, an anti-pattern that does not harden the integration.

  10. Q10D1 · Applications and IntegrationSelect one

    A 180-page contract PDF is reused across dozens of extraction requests per day. Base64 bytes are re-sent every call, dominating bandwidth and latency. What is the MOST cost-effective input method?

    • A. Upload the PDF once via the Files API and reference it by file_id in each request.
    • B. Extract the text locally and paste it into every prompt.
    • C. Convert each page to an image and send images every call.
    • D. Keep sending base64 but enable streaming.
    Show answer

    Answer: A.

    The Files API uploads once and references by file_id, eliminating repeated uploads. A is correct. B still re-sends large text each call. C is even heavier (images per page every time). D (streaming) affects delivery, not the repeated upload cost.

  11. Q11D1 · Applications and IntegrationSelect one

    A caching layer marks a 12,000-token prefix with cache_control, but cache hit rate is near zero. Inspection shows the system prompt embeds the current UTC timestamp. Which statement is correct?

    • A. Any byte change in the prefix (the timestamp) invalidates the cache; move per-request values after the cache breakpoint so the prefix stays byte-identical.
    • B. Caching is disabled on prefixes over 10,000 tokens.
    • C. The prefix is too short to cache; pad it to 16,000 tokens.
    • D. Switch the TTL to 1 hour and the timestamp will be ignored.
    Show answer

    Answer: A.

    A cache hit requires a byte-identical prefix; a per-request timestamp changes it every call. A fixes it by keeping variable content after the breakpoint. B invents a size cap. C misdiagnoses (12k is above the ~1024 minimum). D is wrong: a longer TTL still requires identical bytes to hit.

  12. Q12D1 · Applications and IntegrationSelect one

    Under sustained load the app hits 429 on the ITPM axis first, even though RPM is fine. Which change addresses the actual limiting axis MOST directly?

    • A. Reduce input tokens per request (cache the shared prefix, trim context) and/or request a higher tier.
    • B. Send more requests per minute since RPM has headroom.
    • C. Increase max_tokens so fewer requests are needed.
    • D. Lower temperature to reduce token usage.
    Show answer

    Answer: A.

    The limit hit is input-tokens-per-minute, so cutting input tokens (caching, trimming) or raising the tier is the direct fix. A is correct. B ignores the binding axis. C raises OUTPUT tokens, worsening a different axis. D (temperature) does not change token counts.

  13. Q13D1 · Applications and IntegrationSelect one

    A developer sends messages beginning with a role: 'assistant' turn and gets 400 invalid_request. What is the rule?

    • A. messages must start with a user turn and alternate; prior assistant blocks are appended verbatim as assistant, tool outputs as user with tool_result blocks.
    • B. messages may start with either role; the 400 is a rate limit.
    • C. The first turn must be system inside messages.
    • D. Assistant turns are only allowed when streaming.
    Show answer

    Answer: A.

    Conversations start with user and alternate. A states the rule. B is wrong (a 400 is not a 429). C is wrong: system is a top-level field, not a message entry on current models. D invents a streaming constraint.

  14. Q14D1 · Applications and IntegrationSelect one

    A support engineer asks for the single most useful thing to include in application logs so Anthropic support can trace a specific failed call. What is it?

    • A. The response request-id (correlated with your own request ID).
    • B. The full API key used for the call.
    • C. The user's raw PII for context.
    • D. The max_tokens value only.
    Show answer

    Answer: A.

    The request-id is the correlation handle support and your traces both key on. A is correct. B logs a secret (security violation). C logs PII (hygiene violation). D is insufficient and not a trace handle.

  15. Q15D1 · Applications and IntegrationSelect one

    A harness on Fable 5.1 trims history by editing and reordering earlier turns; users report errors about invalid thinking blocks. Which redesign is correct?

    • A. Make the harness append-only: freeze system/tools, express mid-session changes as role: 'system' messages where supported, and trim only server-side via context editing/compaction.
    • B. Keep mutating history; the error is transient, so add backoff.
    • C. Disable thinking on Fable 5.1 to allow edits.
    • D. Lower temperature so thinking blocks stop appearing.
    Show answer

    Answer: A.

    Fable 5.1 invalidates later thinking when earlier turns change, so the harness must be append-only and trim server-side. A is correct. B blind-retries a deterministic design error. C is impossible (Fable 5.1 always thinks). D does not stop thinking blocks from being produced.

  16. Q16D1 · Applications and IntegrationSelect two

    A migration keeps the same Messages API code but must run through Amazon Bedrock for compliance. Which TWO things actually change versus the direct Anthropic API?

    • A. The client constructor (AnthropicBedrock) and authentication (AWS IAM / SigV4).
    • B. The model ID format (Bedrock-style model identifiers).
    • C. The fundamental request shape of messages/system/max_tokens.
    • D. The meaning of stop_reason values.
    • E. Whether max_tokens is required.
    Show answer

    Answer: A and B.

    Across Bedrock/Vertex the request shape and stop_reason semantics are the same; what changes is the client/auth and the model ID format. A and B are correct. C, D and E are unchanged — the whole point is portability of the request body and control fields.

  17. Q17D1 · Applications and IntegrationSelect one

    A code path calls client.messages.create with max_tokens omitted, expecting a sensible default, and receives 400. Which is the correct explanation?

    • A. max_tokens is required and caps generated output; there is no implicit default, so an omitted value is an invalid request.
    • B. max_tokens is optional; the 400 must be a transient overload.
    • C. max_tokens sets the context window and must equal it.
    • D. max_tokens is only required when tools are declared.
    Show answer

    Answer: A.

    max_tokens is a required field capping generated output. A is correct. B misreads a deterministic 400 as transient. C confuses it with the context window. D invents a tools-only condition.

  18. Q18D1 · Applications and IntegrationSelect one

    An SDK client is created once and shared across an async web server handling thousands of concurrent requests. Which configuration best balances reliability and correctness?

    • A. Use the async client, set a deliberate timeout, enable bounded SDK retries for transient errors, and cap concurrency to respect rate limits.
    • B. Create a new synchronous client per request inside the event loop with no timeout.
    • C. Disable all retries and timeouts to maximise throughput.
    • D. Use unbounded concurrency so all requests fire at once.
    Show answer

    Answer: A.

    An async client with a timeout, bounded retries, and a concurrency cap is the robust pattern. A is correct. B blocks the loop with a sync client and no timeout. C removes safety nets for transient errors. D ignores rate limits and causes 429 storms.

  19. Q19D2 · Model Selection and OptimizationSelect one

    A pipeline runs 200,000 short sentiment classifications overnight; latency is irrelevant and the label set is simple. Which choice is MOST cost-effective?

    • A. Haiku 4.5 via the Message Batches API (cheapest tier + 50% batch discount).
    • B. Opus 5 synchronously with high concurrency for accuracy.
    • C. Sonnet 5 with streaming to speed each call.
    • D. Fable 5.1 with xhigh effort for the best labels.
    Show answer

    Answer: A.

    Simple, high-volume, latency-tolerant work maps to the cheapest capable model plus batching. A is correct. B and D are over-engineered and costly for trivial labels. C (streaming) does not reduce token cost and is irrelevant to an offline job.

  20. Q20D2 · Model Selection and OptimizationSelect one

    A team wants a fixed reasoning budget of 1,500 thinking tokens. On which current model can they set budget_tokens, and what should the others use?

    • A. Only Haiku 4.5 accepts budget_tokens; Opus 5, Sonnet 5 and Fable 5.1 use thinking: {type: 'adaptive'} with effort levels.
    • B. All current models accept budget_tokens if under 2,048.
    • C. Only Opus 5 accepts budget_tokens; the rest have no thinking.
    • D. budget_tokens must be paired with top_p on every model.
    Show answer

    Answer: A.

    budget_tokens is Haiku-4.5-only; other current models use adaptive thinking with effort. A is correct. B is false (it returns 400 elsewhere). C is false (the others do support thinking, just adaptive). D invents a pairing requirement.

  21. Q21D2 · Model Selection and OptimizationSelect one

    A task sends 4,000 input + 800 output tokens, 30,000 times/day. Sonnet 5 ($2/$10 per MTok) and Haiku 4.5 ($1/$5) both clear the quality bar. What is the daily cost difference, and which is cheaper?

    • A. Haiku 4.5 (~$240/day) is cheaper than Sonnet 5 (~$480/day) by about $240/day.
    • B. They cost the same because token counts are identical.
    • C. Sonnet 5 is cheaper because larger models batch better.
    • D. Haiku 4.5 is ~$480/day and Sonnet 5 ~$240/day.
    Show answer

    Answer: A.

    Haiku: input 4000*30000*$1/1e6 = $120 + output 800*30000*$5/1e6 = $120 = $240/day. Sonnet: $240 input + $240 output = $480/day. A is correct (~$240 difference). B ignores that per-token prices differ. C invents a batching effect. D reverses the arithmetic.

  22. Q22D2 · Model Selection and OptimizationSelect one

    Most traffic is trivial; a small tail needs deep reasoning; the hard cases must not degrade and cost must stay low. A junior proposes escalating whenever Claude self-reports confidence below 0.8. What is the correct design?

    • A. Cascade Haiku 4.5 → Sonnet 5 → Opus 5, escalating on a programmatic signal (schema-validation failure, refusal, or a classifier label), not self-reported confidence.
    • B. Run everything on Opus 5 to be safe.
    • C. Run everything on Haiku 4.5 to be cheap.
    • D. Escalate on the confidence score exactly as proposed.
    Show answer

    Answer: A.

    A cascade with a programmatic escalation gate keeps the bulk cheap without hurting the hard tail. A is correct. B is over-engineered/costly; C fails the hard cases (constraint-blind); D relies on self-reported confidence (anti-pattern #4).

  23. Q23D2 · Model Selection and OptimizationSelect two

    Users say a chat UI 'feels sluggish' although total task time is acceptable, and output quality must not change. Which TWO levers best improve the perceived experience?

    • A. Enable streaming so tokens render as they are produced (better time-to-first-token).
    • B. Cache the stable prefix so the model starts generating sooner.
    • C. Switch to Opus 5 for every request.
    • D. Raise max_tokens to 8,000.
    • E. Set effort to xhigh.
    Show answer

    Answer: A and B.

    Perceived latency is time-to-first-token; streaming plus prefix caching (skipping prefix reprocessing) both get output flowing sooner without changing quality. A and B are correct. C, D and E all add latency and/or cost and change behaviour.

  24. Q24D2 · Model Selection and OptimizationSelect one

    For a fair offline comparison of two prompt versions, a developer runs each once at temperature: 0 on a floating model alias and treats identical scores as proof of equivalence. What is the flaw?

    • A. temperature: 0 reduces variance but is not byte-deterministic, and a floating alias can change under you; pin a snapshot and run multiple trials per version.
    • B. There is no flaw; temperature 0 guarantees identical output.
    • C. The flaw is not using xhigh effort for both.
    • D. The flaw is not enabling streaming for reproducibility.
    Show answer

    Answer: A.

    Temperature 0 lowers but does not eliminate variance, and a floating alias undermines reproducibility. A is correct. B is the misconception being tested. C (effort) adds variability, not rigour. D (streaming) is a delivery mechanism unrelated to reproducibility.

  25. Q25D2 · Model Selection and OptimizationSelect two

    Which TWO statements about context windows and output caps across the current lineup are correct?

    • A. Sonnet 5 and Opus 5 each have a 1M context window.
    • B. Haiku 4.5 has a 200k context window.
    • C. Haiku 4.5 has a 1M context window.
    • D. Fable 5.1 has a 128k context window.
    • E. Sonnet 5 has a 200k context window.
    Show answer

    Answer: A and B.

    Sonnet 5 and Opus 5 are 1M context; Haiku 4.5 is 200k. A and B are correct. C is wrong (Haiku is 200k). D confuses Fable 5.1's 128k MAX OUTPUT with its 1M context. E understates Sonnet 5's 1M window.

  26. Q26D2 · Model Selection and OptimizationSelect one

    A 10,000-document extraction runs overnight and is cost-sensitive; Sonnet 5 clears the quality bar. Each request has a 2,000-token stable prefix and 1,000 unique tokens plus 500 output tokens. Which combination is MOST cost-effective?

    • A. Sonnet 5 via the Batches API with cache_control on the 2,000-token shared prefix.
    • B. Opus 5 synchronously, no caching, for maximum quality.
    • C. Sonnet 5 synchronously with caching only (no batching).
    • D. Haiku 4.5 synchronously with xhigh effort.
    Show answer

    Answer: A.

    A latency-tolerant job should stack all three levers: cheapest model that clears the bar, Batches (50% off), and caching the stable prefix. A is correct. B is over-engineered. C leaves the 50% batch discount unused. D is invalid because Haiku 4.5 has no effort parameter.

  27. Q27D2 · Model Selection and OptimizationSelect one

    A team migrating from Sonnet 5 to Opus 5 wants to minimise risk. Which sequence is correct?

    • A. Change the single env-driven model constant behind a flag, re-run the golden eval set, check for removed params and cost/latency deltas, then roll out to a small slice before ramping.
    • B. Swap the model ID in every file and deploy straight to production.
    • C. Let each microservice choose its own model at runtime.
    • D. Change only in production to get real signal fastest.
    Show answer

    Answer: A.

    Centralised, flagged, eval-gated, gradually ramped migration is the safe path. A is correct. B risks a wide blast radius; C causes model drift and inconsistency; D changes production with no safety net.

  28. Q28D3 · Agents and WorkflowsSelect one

    A workflow has three fixed, ordered stages (transcribe → summarise → classify), each with known inputs and outputs. A developer proposes an autonomous agent that decides the steps. What is the BEST design and why?

    • A. Prompt chaining: predictable, testable and cheap for fixed sequential subtasks; reserve agents for paths that cannot be predetermined.
    • B. An autonomous agent, because agents are always more capable.
    • C. A single mega-prompt that does all three at once.
    • D. Orchestrator-workers with dynamic subagent spawning.
    Show answer

    Answer: A.

    Fixed ordered steps are the definition of prompt chaining. A is correct. B is over-engineered — autonomy adds unpredictability with no benefit here. C sacrifices control and testability. D is for tasks whose subtasks are unknown at design time, not this one.

  29. Q29D3 · Agents and WorkflowsSelect one

    An agent loop occasionally never terminates. The current code stops when the assistant text contains 'done' and also caps at 5 iterations. What is the correct PRIMARY termination?

    • A. Continue while stop_reason == 'tool_use' and stop on end_turn; keep the iteration cap only as a safety backstop and handle max_tokens/pause_turn/refusal explicitly.
    • B. Rely on the 5-iteration cap as the main stop.
    • C. Expand the keyword list to catch more 'done'-like phrases.
    • D. Lower temperature so the wording is consistent.
    Show answer

    Answer: A.

    Termination must be driven by stop_reason; the cap is a backstop. A is correct. B makes the cap primary (anti-pattern #2). C parses prose (anti-pattern #1). D does not create a reliable control signal.

  30. Q30D3 · Agents and WorkflowsSelect one

    A research agent's context fills with verbose intermediate tool output, degrading later reasoning and inflating cost. Which design fixes this MOST directly?

    • A. Delegate subtasks to subagents with isolated context that return only distilled summaries to the coordinator.
    • B. Increase max_tokens so more output fits.
    • C. Lower temperature to make output shorter.
    • D. Remove all tools from the agent.
    Show answer

    Answer: A.

    Isolated-context subagents keep noisy intermediate work out of the coordinator's window and return only summaries. A is correct. B caps output length, not context bloat. C does not shrink tool output. D removes needed capability (constraint-blind).

  31. Q31D3 · Agents and WorkflowsSelect two

    In the Claude Agent SDK, which TWO options directly enforce least privilege and deterministic safety for an agent that must never force-push or delete data?

    • A. An allowed_tools allowlist restricting the agent to the tools it genuinely needs.
    • B. A blocking PreToolUse hook that denies matching destructive commands.
    • C. Setting max_turns to 1000.
    • D. A longer, sterner system_prompt forbidding the actions.
    • E. permission_mode: 'bypassPermissions'.
    Show answer

    Answer: A and B.

    An allowlist plus a blocking hook enforce least privilege and hard stops deterministically. A and B are correct. C weakens the backstop. D is prompt-as-enforcement (anti-pattern #3). E removes safeguards entirely.

  32. Q32D3 · Agents and WorkflowsSelect one

    A team wants Anthropic to host the agent loop and sandbox so they can ship standard agentic tasks fast with minimal ops. Which option fits, and what is the main trade-off?

    • A. Managed Agents — Anthropic hosts loop and sandbox; the trade-off is less control over the environment and tools.
    • B. Self-hosted Agent SDK — Anthropic still runs the sandbox for you.
    • C. A raw Messages API loop with no tools — Anthropic manages termination.
    • D. A cron job invoking a single prompt — this is a managed agent.
    Show answer

    Answer: A.

    Managed Agents host the loop and sandbox with lower environment control. A is correct. B misstates the SDK (you host it). C provides no managed loop/sandbox and you own termination. D is not an agent at all.

  33. Q33D3 · Agents and WorkflowsSelect one

    A single agent is configured with 18 tools and frequently selects the wrong one. Which remedy aligns with Anthropic guidance?

    • A. Reduce to ~4–5 focused tools, split responsibilities across subagents, and use the tool search tool with defer_loading for the larger catalogue.
    • B. Add more finely described tools to cover every edge case.
    • C. Force a specific tool on every turn with tool_choice.
    • D. Raise max_turns so it eventually picks correctly.
    Show answer

    Answer: A.

    Too many tools is anti-pattern #8; the fix is fewer active tools, subagent specialisation, and tool search + defer_loading. A is correct. B worsens overload. C is brittle and breaks on Fable 5.1. D does not improve selection.

  34. Q34D3 · Agents and WorkflowsSelect one

    An agent loop returns stop_reason == 'max_tokens' mid-answer, and the harness marks the turn complete. What is the correct handling?

    • A. Recognise the output was truncated: continue the response (or raise max_tokens), not mark it done.
    • B. Treat max_tokens as end_turn; the answer is complete.
    • C. Retry the whole request at temperature: 0.
    • D. Switch to Haiku 4.5 and resend.
    Show answer

    Answer: A.

    max_tokens means truncation, not completion. A recovers the missing content. B silently truncates work (silent-failure). C regenerates rather than continuing and loses the partial. D changes the model without recovering truncated content.

  35. Q35D3 · Agents and WorkflowsSelect one

    A task must run several independent lookups and then aggregate them; the results do not depend on each other. Which pattern fits BEST?

    • A. Parallelization: run the independent subtasks concurrently and aggregate their results.
    • B. Prompt chaining, feeding each result into the next.
    • C. Evaluator-optimizer with a critique loop.
    • D. Routing to a single specialised path.
    Show answer

    Answer: A.

    Independent concurrent subtasks with aggregation is parallelization. A is correct. B serialises independent work needlessly. C is for iterative quality improvement. D dispatches one category, not many concurrent subtasks.

  36. Q36D3 · Agents and WorkflowsSelect two

    A research coordinator delegates topics to worker subagents. Which TWO choices keep the coordinator's context clean and costs down?

    • A. Give each worker its own isolated context window containing only its task slice.
    • B. Have each worker return a short distilled summary to the coordinator.
    • C. Return each worker's full transcript, including intermediate tool output.
    • D. Run all work in the coordinator's single shared context.
    • E. Give every worker all 18 tools.
    Show answer

    Answer: A and B.

    Isolated worker contexts plus distilled-summary returns keep the coordinator window clean and cheap. A and B are correct. C and D reintroduce bloat; E is anti-pattern #8 (too many tools).

  37. Q37D4 · Prompt and Context EngineeringSelect one

    An endpoint on Sonnet 5 must return schema-conforming JSON on every call. Which approach is MOST reliable?

    • A. output_config.format with a JSON schema (or a strict: true tool schema) plus a validation-retry loop that feeds the error back.
    • B. Ask for JSON in the prompt and set temperature: 0.
    • C. Raise max_tokens so the JSON always fits.
    • D. Prefill { and trust the output without validation.
    Show answer

    Answer: A.

    Schema-constrained structured output plus validation-retry is the reliable pattern. A is correct. B is prompt-only and drifts. C addresses truncation, not conformance. D omits validation (silent-failure risk).

  38. Q38D4 · Prompt and Context EngineeringSelect one

    A validation-retry loop rejects the model's JSON but re-sends only the original prompt, so the same error recurs until attempts run out. What is the fix?

    • A. Append the failing output and the specific error (e.g. 'total must be a number') as a new user turn so the next attempt targets that exact defect.
    • B. Increase the retry count to 20 and keep sending the same prompt.
    • C. Switch to Opus 5 for the retries.
    • D. Lower max_tokens on each retry.
    Show answer

    Answer: A.

    Effective validation-retry feeds back the concrete error so the model can correct it. A is correct. B blindly repeats the same mistake (recall-only). C changes the model without telling it what was wrong. D can worsen truncation and conveys nothing.

  39. Q39D4 · Prompt and Context EngineeringSelect two

    A long-running agent slows and loses focus as history grows. The narrative must stay coherent, but old tool results are large and stale. Which TWO server-side controls address this correctly?

    • A. Context editing to clear the stale, no-longer-needed tool results.
    • B. Compaction to summarise older turns while preserving the narrative.
    • C. Increasing max_tokens.
    • D. Adding more tools to help it focus.
    • E. Raising temperature.
    Show answer

    Answer: A and B.

    Context editing clears stale tool outputs; compaction condenses the narrative while keeping the thread. A and B are correct. C caps output length only. D and E do nothing for bloat/drift.

  40. Q40D4 · Prompt and Context EngineeringSelect one

    A long-document QA prompt places the user's question first, then 200 pages of source, and answers are poorly grounded while caching never helps. Which single reordering fixes both problems?

    • A. Put stable instructions and the documents first (with a cache breakpoint) and the question last.
    • B. Move the question into the system prompt but keep it before the documents.
    • C. Interleave the question between every page.
    • D. Split the documents across many random requests.
    Show answer

    Answer: A.

    Documents-first, query-last improves grounding and makes the stable prefix cacheable in one change. A is correct. B still precedes the evidence with the query. C and D harm coherence and caching.

  41. Q41D4 · Prompt and Context EngineeringSelect one

    A team needs reliable structured extraction on Fable 5.1, but forcing a specific tool with tool_choice returns 400. What should they use instead?

    • A. output_config.format with a JSON schema (or strict: true, or auto + a clear instruction), since Fable 5.1 rejects forced tool choice.
    • B. Force the tool anyway and retry the 400 with backoff.
    • C. Disable thinking on Fable 5.1.
    • D. Lower max_tokens until the tool is chosen.
    Show answer

    Answer: A.

    Fable 5.1 rejects any/forced tool choice, so structured outputs or strict/auto are the paths. A is correct. B retries a deterministic 400. C is impossible and irrelevant. D does not influence tool choice.

  42. Q42D4 · Prompt and Context EngineeringSelect two

    A summariser handles common documents well but repeatedly mishandles a rare handwritten-form format. Which TWO prompt changes best improve it?

    • A. Add representative few-shot examples that include the handwritten-form edge case.
    • B. Increase example diversity to cover the range of formats.
    • C. Add ten more examples of the common format.
    • D. Raise temperature for more variety.
    • E. Remove all examples so the model generalises.
    Show answer

    Answer: A and B.

    Edge-case failures call for representative edge-case examples and greater diversity. A and B are correct. C adds more of the already-handled case (recall-only). D adds variability. E discards the steering entirely.

  43. Q43D5 · Tools and MCPsSelect one

    Claude frequently calls the wrong tool and passes malformed arguments for a closed-set field like priority. What is the FIRST and MOST effective fix?

    • A. Rewrite the tool description (what/when/when-not/returns/side-effects) and tighten input_schema with types, an enum for the closed set, and required.
    • B. Lower temperature.
    • C. Switch to Opus 5.
    • D. Increase max_tokens.
    Show answer

    Answer: A.

    The description and schema are the primary levers for correct tool use; an enum prevents invented values. A is correct. B, C and D are secondary and do not address a vague description or loose schema.

  44. Q44D5 · Tools and MCPsSelect two

    A get_order tool times out during a parallel tool call alongside a successful search_kb. Which TWO actions keep the protocol valid and let Claude recover?

    • A. Return a tool_result for the failed call with the same tool_use_id, an error message, and is_error: true.
    • B. Return the successful search_kb result in the same user turn, keyed to its tool_use_id.
    • C. Omit the failed tool's result so Claude ignores it.
    • D. Merge both outputs into one tool_result.
    • E. Raise an exception and end the conversation.
    Show answer

    Answer: A and B.

    Every tool_use needs a matching tool_result in the same turn; errors use is_error: true. A and B are correct. C causes a 400 and hides the failure (silent-failure). D breaks id matching. E abandons the run and loses diagnostics.

  45. Q45D5 · Tools and MCPsSelect two

    Which TWO capabilities are server-side (Anthropic-hosted) built-in tools rather than client-side custom tools?

    • A. Web search.
    • B. Code execution in a sandbox.
    • C. Your internal orders database query function.
    • D. A local shell script you maintain.
    • E. Your company's billing API.
    Show answer

    Answer: A and B.

    Web search and code execution are Anthropic-hosted server-side tools. A and B are correct. C, D and E run in your code/infrastructure and are client-side custom tools.

  46. Q46D5 · Tools and MCPsSelect one

    A team wants the SAME tools available in Claude Code, Claude Desktop and their production Messages API service, maintained in one place, with a remote server secured properly. What is the best mechanism and its security requirement?

    • A. Build an MCP server and connect each host to it (stdio/Streamable HTTP); secure the remote server with OAuth 2.1.
    • B. Copy the custom tool JSON into each application.
    • C. Write a Skill in .claude/skills/ and share the folder.
    • D. Put the tool definitions in CLAUDE.md.
    Show answer

    Answer: A.

    MCP is the cross-host reuse mechanism; remote servers use OAuth 2.1. A is correct. B duplicates maintenance. C is Claude-Code-only packaged know-how, not a cross-host tool server. D holds instructions, not executable tools.

  47. Q47D5 · Tools and MCPsSelect one

    In MCP, which primitive is model-controlled (Claude decides when to invoke it), and how do the other two differ?

    • A. Tools are model-controlled; resources are application-controlled; prompts are user-controlled.
    • B. Resources are model-controlled; tools are user-controlled; prompts are application-controlled.
    • C. Prompts are model-controlled; tools and resources are transports.
    • D. All three are model-controlled.
    Show answer

    Answer: A.

    Tools (model), resources (application), prompts (user) is the correct control mapping. A is correct. B and C scramble the mapping; transports (stdio/HTTP) are not a primitive. D is false — only tools are model-controlled.

  48. Q48D6 · Security and SafetySelect one

    An agent summarises fetched web pages. One page hides an HTML comment instructing Claude to call transfer_funds, and the agent nearly complies. Which combination BEST prevents the transfer?

    • A. Wrap fetched content in content boundaries as data AND enforce a PreToolUse hook that blocks transfer_funds without human approval.
    • B. Add 'never trust web pages' to the system prompt and rely on the model.
    • C. Lower temperature and retry the summary.
    • D. Switch to a larger model that resists injection.
    Show answer

    Answer: A.

    This is indirect injection; the defence is content boundaries (treat external text as data) plus a deterministic hook on the dangerous tool. A is correct. B is prompt-as-enforcement (anti-pattern #3). C and D do not stop injection — temperature and model size are not injection defences.

  49. Q49D6 · Security and SafetySelect two

    A read-only support agent currently has get_order, search_kb, issue_refund and delete_account tools. Which TWO changes best satisfy least privilege and secrets hygiene?

    • A. Remove issue_refund and delete_account from the agent's allowlist; route those through a separate human-approved workflow.
    • B. Keep the API key in environment variables / a secret manager, never in prompts or CLAUDE.md.
    • C. Keep all tools but add logging of refunds and deletes to catch misuse.
    • D. Document the API key in CLAUDE.md for the team.
    • E. Add a system-prompt rule telling the agent not to use refund/delete.
    Show answer

    Answer: A and B.

    Least privilege removes unneeded powerful tools; secrets live in env/secret manager. A and B are correct. C only detects harm after the fact; D leaks the key into version control; E is prompt-as-enforcement (anti-pattern #3).

  50. Q50D6 · Security and SafetySelect two

    A healthcare app must process PHI with zero data retention and FedRAMP High. Which TWO decisions fit the constraints?

    • A. Access Claude via Bedrock or Vertex under a BAA / FedRAMP High authorisation.
    • B. Choose a ZDR-eligible model rather than Fable 5.1, and redact PHI to the minimum needed.
    • C. Use Fable 5.1 because it is the most capable model.
    • D. Put PHI in the system prompt so Claude always has context.
    • E. Disable all logging and hooks to reduce data footprint.
    Show answer

    Answer: A and B.

    Bedrock/Vertex provide FedRAMP High and BAA coverage; a ZDR-eligible model with PHI minimisation meets zero-retention. A and B are correct. C breaks ZDR (Fable 5.1 mandates 30-day retention). D over-shares PHI. E removes enforcement/auditability.

  51. Q51D6 · Security and SafetySelect one

    In a layered-guardrail design (detect → instruct → enforce → verify → approve), where must 'no production deletes without approval' actually be enforced?

    • A. The enforce layer — a deterministic PreToolUse hook that blocks the delete (exit 2) and routes to human approval.
    • B. The instruct layer — a strong sentence in the system prompt.
    • C. The detect layer — an input classifier.
    • D. The verify layer — output validation after the fact.
    Show answer

    Answer: A.

    Critical, irreversible rules must live in the enforce layer as deterministic hooks. A is correct. B is prompt-as-enforcement (anti-pattern #3). C detects but does not block. D checks after the action has already run.

  52. Q52D7 · Claude CodeSelect two

    A team must block the rm -rf command family for everyone in Claude Code, in a way no developer can override, while still running non-interactively in CI. Which TWO steps achieve this?

    • A. Add Bash(rm -rf:*) to permissions.deny and enforce it via managed policy so it is non-overridable (deny beats allow).
    • B. Run in CI headless with claude -p ... --output-format json and an explicit --allowedTools allowlist.
    • C. Write 'never run rm -rf' in ./CLAUDE.md.
    • D. Use --permission-mode bypassPermissions in CI for speed.
    • E. Add a PostToolUse hook that apologises after the command runs.
    Show answer

    Answer: A and B.

    A managed deny pattern blocks the family non-overridably, and headless -p with a tool allowlist runs safely in CI. A and B are correct. C is prose guidance, not enforcement. D removes safeguards. E fires after the damage is done.

  53. Q53D8 · Eval, Testing, and DebuggingSelect two

    A model scores 93% overall on evals yet a production incident involves handwritten forms, and the eval pipeline had the generator grade its own answers in the same session. Which TWO practices would have caught and correctly judged this?

    • A. Report per-segment metrics (by document type) and gate CI on the worst segment, not the aggregate.
    • B. Use an LLM-as-judge in a separate session, ideally a different model, against a rubric.
    • C. Have the same session grade its own output for efficiency.
    • D. Track only the aggregate accuracy.
    • E. Skip evals in CI and rely on production monitoring.
    Show answer

    Answer: A and B.

    Per-segment gating surfaces the handwritten-forms failure (anti-pattern #10), and a separate-session/different-model judge avoids self-review bias (anti-pattern #9). A and B are correct. C is same-session self-review (#9); D is aggregate-masking (#10); E removes the safety net entirely.

Last updated Sep 18, 2026