AI Cert Prep
Type to search documentation.

Domains

D1 · Solution Design & Architecture

Translating business problems into Claude solutions, end-to-end reference architecture, choosing between workflow / agentic / augmented-LLM patterns, cloud placement, capacity planning, HA and fallback design, and ADRs.

This domain is worth 17% – roughly 11 of 63 items – and it sets the frame for the whole exam. It tests whether you can take an ambiguous business problem, extract the constraints and success criteria, and produce a defensible end-to-end architecture. Items here rarely have a single “feature” answer; they reward the option that starts from business value and reasons under constraint. The recurring failure mode is over-engineering (reaching for multi-agent orchestration when a linear workflow meets the SLA at lower cost and higher reliability).

Learning objectives

By the end of this page you should be able to:

  1. Run discovery: extract the real problem, success criteria, and hard constraints from a stakeholder brief.
  2. Draw an end-to-end reference architecture (input → processing → output → feedback loop) and name each component’s responsibility.
  3. Choose among workflow, agentic and augmented-LLM patterns using signal-based decision rules.
  4. Decompose work into multi-agent orchestration only when the signals justify it.
  5. Align a design to the business value pillars (efficiency, cost, performance SLAs) and make a build-vs-buy call.
  6. Choose cloud placement (direct API vs Bedrock vs Vertex vs Foundry) on data-residency, IAM and commitment grounds.
  7. Do capacity planning against rate-limit tiers and design HA and fallback.
  8. Record decisions as ADRs.

1.1 Discovery: from business problem to solvable spec

Architects are handed symptoms, not specs. Discovery converts “we want to use AI for support” into a bounded problem with measurable success criteria and enumerated constraints. If you skip this, every downstream decision is unanchored.

The discovery question set

CategoryQuestions to askWhy it changes the design
Problem & valueWhat decision or task are we automating? What is the value pillar – efficiency, cost, or an SLA?Determines whether latency, accuracy or unit cost dominates
Success criteriaWhat does “good enough” look like numerically (accuracy, p95 latency, deflection rate, cost per task)?Becomes the eval target and the go/no-go gate
Volume & shapeRequests per second, peak vs average, sync vs async, payload sizesDrives capacity planning and Batch vs real-time
DataSensitivity class, residency, retention obligations, sources, freshnessDrives cloud placement, RAG design, compliance
ConstraintsExisting cloud commitments, IAM, budget ceiling, deadline, team skillsNarrows build-vs-buy and placement
Risk & reversibilityWhat is the blast radius of a wrong answer? Which actions are irreversible?Determines human gates and guardrail depth
IntegrationWhat systems must it read/write? What identity model?Drives tool design, MCP vs API, authn/authz

Exam signal

When a stem gives you a vague goal plus one hard number (a latency SLA, a monthly budget, a residency requirement), the correct answer is the one that treats that number as the binding constraint and rejects options that ignore it. “Constraint-blind” is the most common wrong-answer pattern in this domain.

Success criteria must be measurable and per-segment

“Improve support quality” is not a success criterion. “Resolve ≥ 60% of Tier-1 billing tickets without escalation, p95 latency under 6 s, cost under $0.05 per resolved ticket, with no regression on refund-related tickets” is. Note the per-segment clause – aggregate targets hide the failures the exam wants you to catch.


1.2 The end-to-end reference architecture

Every Claude solution decomposes into four stages plus cross-cutting concerns. Learn to place any component into this skeleton.

text
INPUT PROCESSING OUTPUT FEEDBACK LOOP
┌───────────────┐ ┌───────────────────────┐ ┌───────────────┐ ┌───────────────────┐
│ user / event │ │ orchestration layer │ │ validation / │ │ evals (offline) │
│ API / webhook │──▶│ ┌─────────────────┐ │──▶│ structured │──▶│ online metrics │
│ queue / batch │ │ │ Claude model(s) │ │ │ output check │ │ traces + cost │
│ document │ │ │ + tools (MCP) │ │ │ human gate │ │ user thumbs / QA │
└───────────────┘ │ │ + retrieval/RAG │ │ └───────┬───────┘ └─────────┬─────────┘
│ └─────────────────┘ │ │ │
└───────────┬───────────┘ ▼ │
│ downstream systems │
cross-cutting: auth · secrets · guardrails · observability · caching ◀─────┘ (retrain prompts,
reroute models,
tune retrieval)
StageResponsibilityKey decisions
InputNormalise and admit workSync API vs webhook vs queue vs Batch API; payload validation
ProcessingReason and actModel choice/routing, tools, retrieval, orchestration pattern
OutputGuarantee shape and safetyStructured outputs, validation-retry, human gate on irreversible actions
FeedbackImprove the systemOffline evals, online metrics, traces, cost telemetry feeding back into prompts/models/retrieval

The feedback loop is not optional at Professional level. A design with no path from production signals back into prompts, routing and retrieval is incomplete – expect distractors that omit it.


1.3 Choosing the pattern: workflow vs agentic vs augmented-LLM

This is the single most tested decision in D1. The default should be the simplest pattern that meets the criteria. Escalate to agentic only when the task genuinely requires dynamic, model-driven control flow.

PatternWhat it isChoose when the stem shows…Avoid when…
Augmented LLMA single call with retrieval + tools + structured outputTask is one bounded step; deterministic inputs; tight latency/cost SLAThe task needs multiple dependent steps with branching
Workflow (orchestrated)Predefined graph of steps; code owns control flow; models fill stepsSteps are known and stable; you can enumerate the DAG; you need reproducibility and easy debuggingThe path genuinely can’t be known ahead of time
AgenticModel decides next action in a loop until stop_reasonOpen-ended goal; number/order of steps unknown; tool use is exploratoryA fixed workflow meets the SLA more cheaply and reliably (most cases)
text
Is the sequence of steps knowable in advance?
├─ Yes → Can it be done in one call?
│ ├─ Yes → AUGMENTED LLM (retrieval + tools + structured output)
│ └─ No → WORKFLOW (orchestrated DAG; code owns control flow)
└─ No → Does the task truly need dynamic decisions AND is the cost/latency/reliability
hit justified by the value?
├─ Yes → AGENTIC (loop on stop_reason; bounded tools; guardrails)
└─ No → Re-decompose into a WORKFLOW

Control-flow anti-patterns (exam favourites)

An agentic loop must terminate on stop_reason (end_turn / tool_use), never by parsing natural language for “done” and never with an arbitrary iteration cap as the primary stop. A cap is a safety backstop, not the control mechanism. Options that use string matching or a hard cap alone are wrong.


1.4 Multi-agent orchestration and decomposition

Multi-agent (coordinator + subagents) is powerful and expensive. Each subagent multiplies token cost and latency and adds failure surface. Decompose into multiple agents only when the signals are present.

Signal for multi-agentSignal against (keep it single/workflow)
Genuinely parallel subtasks with independent contextSteps are sequential and share context
Distinct skill/tool sets that would bloat one agent’s tool listA single tool set under ~5–7 tools covers it
Need for context isolation (a subagent explores without polluting the coordinator)Latency/cost budget is tight
Long-horizon research/synthesis where fan-out/fan-in helpsReproducibility and easy debugging are priorities
text
┌──────────────┐
│ Coordinator │ owns plan, aggregates, applies guardrails
└──────┬───────┘
┌─────────┼─────────┐
▼ ▼ ▼
┌───────┐ ┌───────┐ ┌───────┐ subagents: isolated context windows,
│ sub A │ │ sub B │ │ sub C │ narrow tool allowlists, own model tier
└───────┘ └───────┘ └───────┘ (e.g. Haiku for retrieval, Opus for synthesis)

Assign cheaper models to narrow subagents (Haiku 4.5 for extraction/routing) and reserve Opus 5 for the coordinator or the hardest synthesis step. This is a portfolio decision, revisited in D2.


1.5 Aligning to business value pillars

Every design serves one dominant pillar; naming it resolves most trade-offs.

PillarDominant metricDesign leversTypical model posture
Efficiency (throughput / deflection)tasks automated, deflection rateworkflow simplification, batching, cachingSonnet 5 default, Haiku for high-volume
Cost (unit economics)cost per taskrouting/cascades, prompt caching, Batch API, output trimmingHaiku 4.5 first, escalate only on failure
Performance SLA (latency)p50/p95 latencyfast mode, smaller model, fewer tool round-trips, streamingHaiku 4.5 / Sonnet 5, fast mode

When two pillars conflict (cheap vs fast, or accurate vs cheap), the stated success criterion breaks the tie. If the stem never states one, the correct answer is usually to go define it with the stakeholder, not to guess.


1.6 Build vs buy

FactorLean buildLean buy
DifferentiationCore to competitive advantageCommodity capability
Team capabilityYou have ML/infra skills to operate itYou lack ops capacity
Time to valueYou have runwayYou need it now
Total cost of ownershipVolume amortises build costLow/uncertain volume
Compliance controlYou need full control of data pathVendor’s certifications suffice

For Claude specifically, “buy” often means using managed agents (Anthropic hosts the loop and sandbox) or a higher-level product, while “build” means the Agent SDK / Tool Runner where you host the loop. Choose managed when you want speed and less ops; choose self-hosted when you need control over the execution environment, data path or custom tooling.


1.7 Cloud placement: direct API vs Bedrock vs Vertex vs Foundry

Claude is available directly and via Amazon Bedrock, Google Vertex AI, and Microsoft Foundry. This is a compliance-and-commitment decision far more than a capability one.

DimensionDirect (Anthropic API)Amazon BedrockGoogle Vertex AIMicrosoft Foundry
Data residencyAnthropic regionsAWS regions incl. FedRAMP HighGCP regions incl. FedRAMP HighAzure regions
IAMAnthropic API keysAWS IAM / SigV4 / rolesGCP IAM / service accountsEntra ID / Azure RBAC
Existing commitmentnoneAWS spend commit / EDPGCP commitAzure commit / MACC
ComplianceAnthropic certsInherit AWS + FedRAMP HighInherit GCP + FedRAMP HighInherit Azure
Latest features firstUsually earliestSlight lagSlight lagSlight lag

Exam signal

“Data must stay in-region”, “FedRAMP High”, “we already have an AWS EDP”, “identity must flow through corporate SSO” → choose the cloud platform whose IAM and residency you already own (Bedrock/Vertex/Foundry). “We want the newest model the day it ships” and no residency constraint → direct API. Don’t pick direct API when a residency or existing-commitment signal is present.


1.8 Capacity planning and rate-limit tiers

Rate limits are enforced per model as RPM (requests/min), ITPM (input tokens/min) and OTPM (output tokens/min), scaled by usage tier. Capacity planning means proving your peak load fits the tier – or designing around it.

  1. Estimate peak: peak_RPM = peak_requests_per_sec × 60; peak_ITPM = peak_RPM × avg_input_tokens; likewise OTPM.

  2. Compare against the model’s tier limits (client.models.retrieve(id) and account tier). If you exceed any of the three, you are limited by that one.

  3. Design around limits: shift latency-tolerant work to the Message Batches API (50% discount, results within 24 h), spread load, request a tier increase, or route overflow to a second model.

  4. Handle 429s with exponential backoff + jitter, honouring the retry-after header. A 429 is expected under burst, not an error to swallow.

python
import time, random, anthropic
client = anthropic.Anthropic()
def call_with_backoff(**kwargs):
for attempt in range(6):
try:
return client.messages.create(**kwargs)
except anthropic.RateLimitError as e:
retry_after = float(getattr(e, "retry_after", 0) or 0)
sleep = retry_after or min(2 ** attempt + random.random(), 30)
time.sleep(sleep)
raise RuntimeError("exhausted retries")

1.9 High availability and fallback design

Production Claude systems must degrade gracefully, not fail hard. Retry 429/500/529 with backoff; for sustained unavailability, fall back to another model or provider.

text
┌─────────────────────────┐
request ──▶ │ primary: claude-opus-5 │──✓──▶ response
└───────────┬─────────────┘
│ 529 overloaded / timeout (after backoff)
▼
┌─────────────────────────┐
│ fallback: claude-sonnet-5│──✓──▶ response (log degraded mode)
└───────────┬─────────────┘
│ still failing
▼
┌─────────────────────────┐
│ cross-provider: Bedrock │──✓──▶ response
│ or queued for retry │
└─────────────────────────┘
FailureMitigation
429 rate limitbackoff + jitter, honour retry-after, spillover routing
529 overloaded / 5xxretry with backoff; fall back to secondary model/provider
Regional outagemulti-region via Bedrock/Vertex; queue and replay
Bad output shapestructured outputs + validation-retry (D4)
Irreversible actionhuman approval gate before execution

Fallback that silently drops capability

Fable 5.1’s thinking blocks are readable only by the producing model or newer; a silent fallback to an older model drops those blocks and can corrupt a multi-turn agentic session. Fallback design must account for feature parity, not just availability. Never mask a degraded path as a normal success.


1.10 Architecture Decision Records (ADRs)

An ADR captures one decision, its context, the options considered, the choice, and the consequences. On the exam, the “best” answer often mirrors ADR discipline: it states the constraint, names the rejected alternative and why, and accepts an explicit trade-off.

markdown
# ADR-014: Retrieval layer for the policy-QA assistant
## Status: Accepted
## Context
Regulated (GDPR); 2M docs; answers must cite source clause; p95 < 8 s; budget $0.04/query.
## Decision
Hybrid retrieval (BM25 + dense) with reranking; Claude Sonnet 5 for synthesis on Vertex AI (EU residency).
## Alternatives considered
- Long-context stuffing (1M): rejected — cost per query > budget, no citation granularity.
- Fine-tuning: rejected — corpus changes weekly; retraining cadence infeasible.
- Direct API: rejected — EU data residency requires Vertex EU region.
## Consequences
+ Citations at clause level; cost within budget.
- Added reranker latency (~300 ms) and an index-refresh pipeline to operate.

1.11 Capacity and cost modelling with arithmetic

Capacity planning at Professional level is a numbers exercise, not a vibe. Prove the load fits the tier and the budget before the design review.

Worked capacity model

A support workload: peak 20 requests/sec, average 8 req/s; each call ~6k input + 1.5k output tokens on Sonnet 5.

text
peak_RPM = 20 req/s × 60 = 1,200 RPM
peak_ITPM = 1,200 × 6,000 = 7,200,000 input tokens/min
peak_OTPM = 1,200 × 1,500 = 1,800,000 output tokens/min

You are limited by whichever of RPM / ITPM / OTPM you breach first. If the tier caps ITPM at 4,000,000, ITPM is the binding limit at ~55% of peak — so ~45% of peak requests must shift to Batch, spill to a second model, or you request a tier increase.

Worked cost model (per month)

Average 8 req/s → 8 × 86,400 = 691,200 req/day ≈ 20.7M req/month.

DesignPer-request costMonthly (20.7M req)
All Sonnet 5 (6k in @ $2, 1.5k out @ $10)6k×$2/1e6 + 1.5k×$10/1e6 = $0.012 + $0.015 = $0.027$559k
+ prompt cache (80% hit on a 4k stable prefix, read 0.1×)saves ~0.8×(4k×$2×0.9)/1e6 = ~$0.0058 → $0.021$435k
Cascade: 70% Haiku ($0.0135), 30% Sonnet ($0.027)0.7×$0.0135 + 0.3×$0.027 = $0.0176$364k
Cascade + cache + 30% batchable at 50% off≈ $0.013$269k

Exam signal

When a stem gives request rate, token sizes and a budget, the correct answer is the design whose arithmetic clears the budget — usually cache + cascade + batch, not “buy a bigger model” or “hope the tier is enough”. Show the maths in your head: cost = (input_tokens × in_price + output_tokens × out_price) ÷ 1e6.

LeverTypical savingPrecondition
Prompt caching40–90% of input cost on the cached prefixStable prefix ≥ ~1024 tokens
Cascade / routing30–70% overallA reliable validation check to escalate on
Batch API50% on eligible trafficLatency tolerance up to 24 h
Output trimming / structured output10–40% output costDon’t drop needed content

1.12 Latency budget decomposition

A p95 SLA is a budget to be allocated across the request path. If the parts sum above the SLA, the design fails before it ships.

text
p95 target: 6,000 ms
├─ network + auth ingress ...... 150 ms
├─ retrieval (hybrid + rerank) .. 400 ms
├─ model TTFT (Sonnet 5) ........ 600 ms
├─ model generation (1.5k tok) .. 3,200 ms
├─ tool round-trip (1 call) ..... 700 ms
├─ output validation ............ 120 ms
└─ headroom ..................... ~830 ms ✓ fits
If the budget is blown by…Lever
Generation timeSmaller/faster model, fast mode, shorter output, streaming (improves perceived latency)
Too many tool round-tripsFewer tools, parallelise independent calls, cache tool results
RetrievalLower k with reranking, warm the index, cache embeddings
Model queueing under loadHigher tier, spillover routing, Batch for non-interactive work

Exam signal

“p95 is over budget and most of it is generation” → shrink the model / output / effort, or stream. Adding a multi-agent layer increases latency; it is the wrong answer whenever a latency budget is the binding constraint.


1.13 Scenario walkthrough: designing an insurance-claims triage assistant

Scenario. A mid-size insurer wants to “use AI to speed up claims”. Discovery surfaces: 12,000 claims/day (peak 3×), each claim has a PDF plus structured metadata; the assistant must classify claim type, extract key fields, flag likely fraud for human review, and draft a customer acknowledgement. Regulated (GDPR, EU residents), on Azure already, budget $40k/month, p95 under 10 s for the interactive draft, and fraud flags must not auto-deny — a human adjuster decides. Historical fraud rate is ~4%.

Expert reasoning trace.

  1. Anchor on constraints. GDPR + EU residents + existing Azure → placement is Microsoft Foundry (Azure) in an EU region; direct API is rejected on residency. Fraud auto-deny is irreversible and regulated → a human gate is mandatory, not optional.

  2. Pattern choice. The steps are enumerable (classify → extract → fraud-score → draft), so this is a workflow, not an agentic loop. Reject multi-agent: the steps are sequential and share context; fan-out buys nothing and adds cost/latency.

  3. Model portfolio. Extraction and classification are narrow, high-volume → Haiku 4.5. Fraud reasoning and the customer draft are higher-stakes → Sonnet 5. Reserve escalation to Opus 5 only for low-confidence fraud cases (a cascade on a validation signal, never on the model’s self-reported confidence).

  4. Capacity + cost. 12k/day base, peak 3× → ~0.4 req/s average, ~1.25 req/s peak — well within tier; no Batch needed for the interactive path, but the nightly bulk re-scoring can use Batch. Cost is dominated by the PDF input tokens → cache the stable system/policy prefix; the arithmetic lands under $40k.

  5. Feedback loop. Adjuster accept/override on fraud flags is the gold-label stream → feeds a per-segment eval (by claim type) and recalibrates the fraud threshold. Without this loop the design is incomplete.

  6. Record the ADR. Placement (Foundry EU), pattern (workflow), portfolio (Haiku/Sonnet/Opus cascade), human gate on fraud, and the rejected alternatives (direct API, multi-agent, auto-deny) with reasons.

Why each tempting alternative is wrong: direct API ignores EU residency; multi-agent over-engineers a sequential workflow; auto-deny removes the mandatory human gate on an irreversible, regulated action; escalating on self-reported confidence is anti-pattern #4; a single aggregate accuracy number would hide per-claim-type failure.


1.14 Common misconceptions

MisconceptionRealityWhy it matters on the exam
“Agentic is more advanced, so it’s the better design.”Agentic is a tool for unknowable control flow; a workflow is better when steps are known.The over-engineering distractor is the most common wrong answer in D1.
“A bigger context window removes the need for RAG.”A 1M window still costs per token and can degrade on needle-in-haystack; retrieval is cheaper and fresher.Distractors offer “stuff the corpus into context” — reject it on cost/freshness.
“The most capable model is the safe default.”The cheapest model that clears the quality bar is the right default; capability is routed to where it’s needed.Blanket-Opus answers blow cost/latency budgets.
“A 429 means something is broken.”429 is expected back-pressure; handle with backoff + retry-after.Answers that treat 429 as fatal or swallow it are wrong.
“Fallback just means retry a cheaper model.”Fallback must preserve feature parity (e.g. Fable 5.1 thinking blocks) and log degraded mode.Silent-downgrade distractor corrupts agentic sessions.
“Cloud placement is a performance choice.”It is primarily a residency/IAM/commitment (compliance) choice.Residency signals in the stem override novelty and cost.
“Success criteria are a single accuracy number.”Criteria must be per-segment and tied to the value pillar.Aggregate-metric trap hides segment failures.

Exam traps in this domain

TrapWhy it is wrong
Reaching for multi-agent orchestration for a task a linear workflow handlesOver-engineered; higher cost, latency and failure surface for no benefit
Terminating an agent loop by parsing text for “done”Should check stop_reason; string matching is brittle and unsafe
Using an iteration cap as the primary stopping mechanismA cap is a backstop; control flow should be driven by stop_reason
Choosing direct API when the stem states EU residency / FedRAMPIgnores a binding constraint; use Bedrock/Vertex/Foundry
Designing with no feedback loopIncomplete architecture; no path from production signals to improvement
Picking the most powerful model everywhere to “be safe”Blows cost/latency budgets; routing/cascades exist for this
Treating a 429 as a fatal errorExpected under burst; retry with backoff + retry-after
Silent fallback to an older model in a Fable 5.1 sessionDrops thinking blocks; corrupts the agentic loop; hides degradation
Setting success criteria as a single aggregate numberMasks per-segment failure; criteria must be per-segment
Skipping discovery and designing from the vague askUnanchored design; the binding constraint is never surfaced
Answering a latency-budget stem by adding a multi-agent layerMulti-agent increases latency; wrong when p95 is binding
“Buy a bigger model” as the cost fix when the arithmetic favours cache + cascade + batchIgnores the cost model; a bigger model raises unit cost
Choosing placement on performance/novelty when a residency signal is presentResidency/IAM/commitment govern placement, not speed
Auto-executing an irreversible/regulated action without a human gateRemoves the mandatory approval gate; unsafe and non-compliant
Treating capacity as “the tier is probably fine” without computing RPM/ITPM/OTPMThe binding limit is whichever of the three you breach first

Practice questions

Q1 · A retailer asks for 'an AI agent to handle customer emails'. Volume is 400 emails/hour, mostly order-status lookups against one API, with a p95 latency target of 5 s and a tight cost ceiling. What should the architect propose FIRST? (Select one)

A. A multi-agent system with a coordinator and specialised subagents for each email type. B. An augmented-LLM or simple workflow: classify intent, call the order-status tool, return a structured reply — because the steps are knowable and the SLA/cost budget favour the simplest pattern. C. An agentic loop with a large tool catalogue so it can handle anything. D. Fine-tuning a model on historical emails before any pipeline exists.

Answer: B. The steps are enumerable (classify → lookup → reply), latency and cost are binding, and the volume is modest. The simplest pattern that meets the criteria wins. Multi-agent (A) and a broad agentic loop (C) are over-engineered, raising cost, latency and failure surface. Fine-tuning (D) is premature with no pipeline or eval baseline.

Q2 · A healthcare provider (EU-based, GDPR, data must not leave the EU) wants a Claude assistant. They already run everything on Google Cloud. Which placement is BEST? (Select one)

A. Direct Anthropic API for earliest access to new models. B. Google Vertex AI in an EU region, inheriting GCP IAM and residency. C. Amazon Bedrock, because it has the most compliance certifications. D. Whichever is cheapest per token.

Answer: B. Two binding constraints — EU residency and an existing GCP footprint (IAM, commitment). Vertex AI in an EU region satisfies both. Direct API (A) risks residency. Bedrock (C) is compliant but ignores the existing GCP investment and identity model. Cost (D) cannot override a legal residency requirement.

Q3 · An agent occasionally never finishes a task; a junior engineer proposes stopping the loop after 10 iterations and also scanning the model's text for the word 'complete'. What is the architect's correct guidance? (Select two)

A. Drive termination from stop_reason (end_turn / no further tool_use). B. Keep the iteration cap, but only as a safety backstop, not the primary control. C. Keep parsing the text for ‘complete’ as the main signal. D. Remove all limits and trust the model to stop. E. Lower the temperature so it stops sooner.

Answer: A and B. Control flow must key off stop_reason; an iteration cap is a legitimate backstop against runaway loops but not the primary mechanism. Parsing natural language (C) is brittle and unsafe. Removing all limits (D) risks runaway cost. Temperature (E) does not govern termination.

Q4 · A stakeholder says 'make support better with AI' and offers no numbers. What is the BEST first action? (Select one)

A. Start building an agentic system immediately. B. Run discovery to define measurable, per-segment success criteria and enumerate constraints before designing. C. Pick Opus 5 because it is the most capable. D. Assume a 90% deflection target and proceed.

Answer: B. With no success criteria or constraints, any design is unanchored. Discovery surfaces the value pillar, numeric targets (per segment), volume, data sensitivity and constraints. Building (A), defaulting to the biggest model (C), or inventing a target (D) all skip the anchoring step the exam rewards.

Q5 · Peak load is 30 requests/sec with ~8k input tokens each. The team hits frequent 429s on the target model's tier. Which combination is the SOUNDEST response? (Select two)

A. Move latency-tolerant jobs to the Message Batches API and add exponential backoff with jitter honouring retry-after. B. Retry immediately in a tight loop until it succeeds. C. Request a higher usage tier and/or route overflow to a second model. D. Swallow the 429 and return an empty result as success. E. Switch every request to Opus 5.

Answer: A and C. ITPM/RPM ceilings are exceeded at peak; shifting tolerant work to Batch (50% discount, 24 h) plus backoff, and raising the tier or spilling over to a second model, are the correct capacity levers. Tight-loop retry (B) worsens the storm. Swallowing errors (D) is a silent-failure anti-pattern. Upgrading every request to Opus (E) raises cost and does not fix rate limits.

Q6 · Which scenario genuinely justifies a multi-agent (coordinator + subagents) design? (Select one)

A. A three-step, sequential document-cleanup pipeline with shared context. B. A research task that fans out into several independent investigations with isolated context, then synthesises results. C. A single order-status lookup with a tight latency budget. D. Any task, to be safe.

Answer: B. Independent parallel subtasks with context isolation and fan-in synthesis are the textbook multi-agent signal. Sequential shared-context work (A) is a workflow. A single lookup (C) is augmented-LLM. “Any task” (D) is the over-engineering trap.

Q7 · A design routes all traffic to Opus 5. Cost is 4× budget but accuracy is fine. What is the BEST optimisation that preserves quality? (Select one)

A. Switch everything to Haiku 4.5 and accept lower accuracy. B. Introduce a cascade: attempt Haiku 4.5 / Sonnet 5 first, escalate to Opus 5 only when confidence/validation checks fail, and add prompt caching for the stable prefix. C. Reduce the number of users. D. Remove the eval suite to save compute.

Answer: B. Cascades route cheap-first and escalate on failure, cutting cost while preserving quality on hard cases; caching amortises the stable prefix. Blanket Haiku (A) sacrifices accuracy. Cutting users (C) or evals (D) does not address unit economics responsibly.

Q8 · The primary model returns 529 (overloaded) under a traffic spike. Which fallback design is BEST? (Select one)

A. Return an error to every user until it recovers. B. Retry with backoff; if still failing, fall back to a secondary model/provider and log the request as served in degraded mode. C. Silently return cached-but-stale answers as if fresh. D. Immediately fail over to a much older model in an active Fable 5.1 thinking session.

Answer: B. Retry-then-fallback with explicit degraded-mode logging keeps the system available and observable. Hard failure (A) is poor HA. Passing stale answers off as fresh (C) is silent failure. Failing an active Fable 5.1 session to an older model (D) drops thinking blocks and corrupts the session.

Q9 · An architect must decide build vs buy for the agent runtime. The company lacks ops capacity, needs it live in six weeks, and the capability is not a differentiator. What is the BEST call? (Select one)

A. Build a custom Agent SDK runtime and sandbox from scratch. B. Use managed agents (Anthropic hosts the loop and sandbox) to reduce ops burden and hit the deadline. C. Delay the project until an ML platform team is hired. D. Buy a competitor’s product with no Claude support.

Answer: B. No ops capacity, tight deadline, commodity capability → buy/managed. Managed agents remove loop and sandbox operations. Building (A) contradicts the constraints; delaying (C) fails the deadline; (D) abandons the requirement.

Q10 · What belongs in the 'feedback loop' stage of a reference architecture? (Select two)

A. Offline eval runs against a golden set on each release. B. Online metrics, traces and cost telemetry that inform prompt/model/retrieval changes. C. The initial request normalisation and payload validation. D. The structured-output schema definition. E. The load balancer in front of the API.

Answer: A and B. The feedback stage closes the loop from production signals back into the system’s prompts, routing and retrieval — offline evals and online telemetry/traces do exactly that. Input normalisation (C) and schema definition (D) belong to input/output stages; the load balancer (E) is infrastructure, not feedback.

Q11 · A stem states a hard p95 latency of 3 s and a modest accuracy bar for a high-volume classification task. Which posture is BEST? (Select one)

A. Opus 5 with xhigh effort for maximum quality. B. Haiku 4.5 (or Sonnet 5) possibly with fast mode, minimal tool round-trips, and prompt caching — because latency and volume are the binding pillars and the accuracy bar is modest. C. A multi-agent pipeline for robustness. D. Long-context stuffing of the entire knowledge base per request.

Answer: B. The binding pillar is latency at volume with only a modest accuracy need — the cheapest/fastest model that clears the bar, with fewer round-trips and caching. Opus + xhigh (A) maximises latency and cost against the SLA. Multi-agent (C) adds latency. Long-context stuffing (D) inflates tokens, cost and latency.

Q12 · Why record an ADR for the workflow-vs-agentic choice? (Select one)

A. To satisfy a documentation quota. B. To capture the context, the rejected alternatives and their reasons, and the accepted trade-off, so the decision can be reviewed and revisited as constraints change. C. Because ADRs replace evals. D. Because agentic systems cannot be built without one.

Answer: B. An ADR makes the reasoning and trade-offs explicit and reviewable — exactly the altitude the exam rewards. It is not a formality (A), does not replace evaluation (C), and is not a technical prerequisite (D).

Q13 · A workload runs 8 req/s average at 6k input + 1.5k output tokens on Sonnet 5 (\$2/\$10 per MTok). Roughly what is the per-request cost, and which lever cuts it MOST for a stable-prefix workload? (Select one)

A. About $0.027; the biggest single lever is prompt caching on the stable prefix. B. About $0.27; switch everyone to Opus 5. C. About $0.003; do nothing. D. Cost is unknowable without a load test.

Answer: A. 6k×\$2/1e6 + 1.5k×\$10/1e6 = \$0.012 + \$0.015 = \$0.027. With a stable prefix, prompt caching (read ≈ 0.1× input) removes most of the input cost — the largest lever here. Opus (B) raises unit cost; $0.003 (C) is off by an order of magnitude; the cost is directly computable (D).

Q14 · A p95 budget of 6 s is being blown, and traces show generation of a long output dominates the time. Which change BEST fits? (Select one)

A. Add a coordinator and three subagents for robustness. B. Use a smaller/faster model or fast mode, shorten/stream the output, and lower effort where adequate — because generation is the dominant term. C. Increase k in retrieval to improve quality. D. Move the interactive path to the Batch API.

Answer: B. The latency budget is dominated by generation, so shrink the model/output/effort or stream. Multi-agent (A) adds latency; larger k (C) adds retrieval time; Batch (D) has up-to-24 h latency and cannot serve an interactive p95 SLA.

Q15 · Peak load is 20 req/s at 6k input tokens each; the tier caps ITPM at 4,000,000. What is the binding constraint and the sound response? (Select two)

A. ITPM is the binding limit (peak ITPM ≈ 7.2M > 4M). B. RPM is the binding limit; nothing else matters. C. Shift latency-tolerant traffic to Batch and/or request a higher tier or spill overflow to a second model. D. Ignore it; bursts are rare. E. Send everything to Opus 5 to be safe.

Answer: A and C. peak_ITPM = 20×60×6,000 = 7.2M, which exceeds the 4M cap, so ITPM binds first. The fix is to reduce interactive ITPM via Batch, a higher tier, or spillover routing. RPM-only (B) misreads the maths; ignoring bursts (D) causes 429 storms; Opus for all (E) raises cost and ITPM.

Q16 · An EU-regulated insurer on Azure wants claims triaged with a fraud flag; fraud must never auto-deny. Which combination is BEST? (Select two)

A. Deploy on Microsoft Foundry in an EU region for residency and existing IAM. B. Route fraud auto-deny straight to the model to save adjuster time. C. Insert a mandatory human approval gate before any denial (irreversible, regulated action). D. Use the direct Anthropic API for the newest model. E. Set a single aggregate accuracy target across all claim types.

Answer: A and C. EU residency + existing Azure → Foundry EU; an irreversible regulated action requires a human gate. Auto-deny (B) removes the mandatory gate; direct API (D) risks residency; a single aggregate target (E) hides per-claim-type failure.

Q17 · A design falls back from Opus 5 to a much older model automatically during a 529 spike, inside an active Fable 5.1 thinking session. Users see corrupted multi-turn behaviour. What is the correct fix? (Select one)

A. Increase the iteration cap. B. Fall back only to a model with feature parity (equal-or-newer thinking support), log the degraded mode, and never silently downgrade a thinking session. C. Disable retries so it fails fast. D. Parse the output text to detect corruption.

Answer: B. Fallback must preserve feature parity — thinking blocks are readable only by the producing model or newer — and degraded mode must be logged, not silent. Iteration caps (A) and text parsing (D) don’t address feature parity; disabling retries (C) hurts availability.

Q18 · A stem gives request rate, token sizes, and a hard monthly budget, and asks the MOST cost-effective design that preserves quality. What is the BEST approach? (Select one)

A. Pick the most capable model so quality is never in question. B. Compute per-request cost, then combine prompt caching on the stable prefix, a cheap-first cascade escalating on validation failure, and Batch for latency-tolerant traffic until the arithmetic clears the budget. C. Guess a design and adjust after launch. D. Remove evaluation to reduce compute cost.

Answer: B. Cost-effectiveness under a stated budget is an arithmetic exercise: cache + cascade + batch layered until cost < budget, with quality preserved on hard cases via the cascade. Blanket-capable (A) overspends; guessing (C) is unanchored; cutting evals (D) removes the quality guard.

Key takeaways

  • Start every design from discovery: the value pillar, measurable per-segment success criteria, and the binding constraints.
  • Decompose into input → processing → output → feedback loop; a design without the loop is incomplete.
  • Choose the simplest pattern that meets the criteria: augmented-LLM → workflow → agentic. Over-engineering is the top wrong answer.
  • Agentic loops terminate on stop_reason; iteration caps are backstops, never the primary control.
  • Multi-agent only when subtasks are genuinely parallel/isolated; assign cheaper models to narrow subagents.
  • Cloud placement follows residency, IAM and existing commitments — Bedrock/Vertex/Foundry for those constraints, direct API otherwise.
  • Plan capacity against RPM/ITPM/OTPM tiers; use Batch API, backoff and spillover routing.
  • Design for failure with retries and model/provider fallback, minding feature parity (Fable 5.1 thinking blocks).
  • Record decisions as ADRs that name rejected alternatives and accepted trade-offs.
  • Model the arithmetic: per-request cost = (in_tok×in_price + out_tok×out_price)/1e6; layer caching + cascade + Batch until it clears the budget.
  • Treat a p95 SLA as a budget decomposed across the request path; if generation dominates, shrink model/output/effort or stream — never add latency with multi-agent.
  • Compute RPM / ITPM / OTPM and design around whichever binds first; “the tier is probably fine” is not capacity planning.
  • Put a human gate on every irreversible/regulated action, and make fallbacks preserve feature parity with degraded-mode logging.

Last updated Sep 18, 2026