AI Cert Prep
Type to search documentation.

Appendix · Claude

Decision Frameworks

Every A-versus-B trade-off the four exams test, consolidated into one reference with the scenario signal that points to each answer.

The exams are trade-off exams. This page consolidates the decisions across all four blueprints. For each: the options, the signal words in a stem, and the exam-correct choice.

Architecture

DecisionOption AOption BSignal → choice
Workflow vs agentWorkflow – predefined code pathsAgent – model directs its own stepsSteps known and repeatable → workflow. Open-ended, path depends on findings → agent. Default to the simplest.
Single call vs chainOne promptPrompt chaining with gatesDistinct stages with checkable intermediate output → chain
RoutingSingle generalist promptClassifier → specialist prompts/modelsHeterogeneous inputs with different handling → route (cheap router)
ParallelizationSequentialSectioning (independent parts) / voting (same task ×N)Independent subtasks or need for consensus → parallel
Orchestrator-workersFixed decompositionOrchestrator decomposes dynamicallySubtasks unknown until inspected → orchestrator
Evaluator-optimizerSingle passGenerate → evaluate → refineClear evaluation criteria and iterative value → loop; evaluator in separate session
HostingManaged Agents – Anthropic hosts loop + sandboxAgent SDK / Tool Runner – you hostNeed control over runtime, network, data locality → self-host; want least ops → managed
Subagents vs dynamic workflowsSubagents/forksScripted dynamic workflowsA few delegated tasks → subagents; dozens–hundreds → dynamic workflows
Build vs escalate (Associate)Configure in claude.ai/ProjectsEscalate to Developer/ArchitectNeeds API, automation, integration, or system-level guarantees → escalate

Loop control and enforcement

DecisionCorrectAnti-pattern
Loop terminationBranch on stop_reasonParse prose for “done”; arbitrary iteration cap as primary stop
Critical business ruleProgrammatic hook (PreToolUse) / tool permissionSystem-prompt sentence
Escalation triggerExplicit request → immediate; beyond capability → resolve-then-escalateSentiment; self-reported confidence
Error reportingStructured: category, retryable, partial resultsGeneric message; silent empty success
Self-reviewFresh session or different modelSame-session “are you sure?”
Quality metricPer-segment / per-document-typeAggregate only
Tool count4–5 focused tools; tool search + defer_loading beyond ~1018 tools loaded upfront

API mechanics

DecisionOption AOption BSignal → choice
Realtime vs BatchMessages APIMessage Batches (50% off, ≤24 h)“Overnight”, “cost matters, latency doesn’t”, bulk → Batch
StreamingNon-streamingSSE streamingInteractive UI, long outputs, perceived latency → stream
Structured outputPrompt “return JSON”output_config.format / strict toolsMachine-consumed output → schema-enforced; always validate + retry
Forced tool calltool_choice: any/toolauto + instruction / strict / structured outputFable 5.1 → B (A is 400)
Thinkingbudget_tokens{"type":"adaptive"}Current models → adaptive (Haiku 4.5 is the exception)
Efforthigh (default)low/medium for mechanical; xhigh for hardestSubagent doing routine work → low/medium
CachingNonecache_control on stable prefixRepeated long prefix ≥1024 tokens → cache; stable first
Context too longClient-side truncationContext editing / compaction (server-side)Verbose tool results → edit; long narrative → compact; Fable 5.1 → server-side only
Persistent stateConversation historyMemory tool / files / DBMust survive compaction or sessions → external
RetryRetry everythingBackoff on 429/5xx/529 only; honour retry-after4xx client errors → fix, don’t retry
Fallback modelAnyNewer-or-equal modelFrom Fable 5.1 to older → thinking blocks dropped; plan for re-planning

Model selection

SignalChoice
Classification, routing, extraction at scale, sub-secondHaiku 4.5
Balanced quality/cost general tasks, customer chatSonnet 5
Complex agentic coding, enterprise defaultOpus 5
Frontier reasoning, budget secondary, accept constraintsFable 5.1
Mixed difficultyCascade cheap → expensive with external validator
Latency criticalSmaller model, streaming, shorter output, caching, fast mode
Cost criticalSmaller model, caching, Batch, trim output, lower effort

Prompting and context

DecisionChoiceSignal
Zero-shot vs few-shotFew-shotFormat-sensitive, ambiguous, edge cases
CoT vs extended thinkingExtended/adaptive thinkingHard reasoning; keep visible output clean
Document placementDocuments first, question lastLong context recall + cacheable prefix
Data vs instruction boundaryXML tagsUntrusted content; injection risk
Prompt as assetVersion, test, cache-stableProduction system prompt
Context isolationSubagent with explicit contextProtect main context; parallel exploration
Plan vs direct (Claude Code)Plan modeMulti-file, architectural, unfamiliar, irreversible
Skill vs CLAUDE.md vs command vs subagent vs MCPSee tableProcedure → Skill; conventions → CLAUDE.md; parameterised one-liner → command; isolated delegated task → subagent; external system → MCP

Tools and MCP

DecisionChoiceSignal
Custom tool vs MCP serverMCPReused across hosts/agents; external system
stdio vs Streamable HTTPHTTPRemote, multi-user, needs OAuth
Tool description qualityRich: what/when/when-not/returnsModel picks wrong tool → fix description first
Destructive tool exposed but unneededRemove itNot “log it” or “confirm it”
IdentityPer-user OAuth propagationMulti-tenant app; authz gap
Parallel tool callsAllowIndependent lookups

RAG (Professional)

DecisionChoiceSignal
RAG vs long context vs fine-tuningRAGLarge/changing corpus, citations needed. Long context: small stable corpus fits 1M. Fine-tuning: style/format, not facts
ChunkingStructural/document-aware; parent-childStructured docs; need precise hits with broad context
RetrievalHybrid (dense + BM25) + rerankerMixed paraphrase and exact-ID queries
Query handlingRewriting / multi-query / HyDEVague or multi-intent questions
Confident-but-wrong after doc refreshSuspect retrieval/indexing firstStale index, chunk drift
Retrieval evalrecall@k, MRR, faithfulness, answer relevanceSeparate retrieval quality from generation quality

Evaluation

DecisionChoiceSignal
GraderExact match → rubric → LLM-as-judge (separate model, calibrated)Increasing subjectivity
A/BPairwise with significanceComparing prompts/models
WhereOffline golden set + online canaryBefore and after release
CIRegression suite gating deployPrompt or model change
Metric viewPer-segment + aggregateAlways both
Root causeIntegration-layer vs model-output vs retrievalCheck traces before changing prompts

Governance and safety

DecisionChoiceSignal
Irreversible / regulated / external / PII actionHuman-in-the-loop gate“must never”, “publish”, “pay”, “delete”, “diagnosis”
Data class → toolFollow classification and policyRestricted data → approved enterprise surface only
RetentionZDR where requiredRegulated; note Fable 5.1 exclusion
Compliance frameworkGDPR (EU personal data), HIPAA (PHI, BAA), FedRAMP High (US federal via Bedrock/Vertex), SOC 2Sector words in the stem
GuardrailsLayeredNever single-layer
DisclosureTell users when AI is involvedExternal comms, decisions about people

Stakeholders and lifecycle (Professional)

DecisionChoiceSignal
DiscoveryStructured interviews → success metrics → constraintsVague requirements
Communicating trade-offsADR + decision matrix + cost model + risk registerNon-technical sponsor
ExpectationsSLA/SLO alignment with measured baselines“Executives expect 100% accuracy”
LifecycleDiscovery → design → build → handoff → monitoring → iteration with exit criteriaHandoff without runbooks is the distractor
Model deprecationTreat as planned lifecycle event with eval-gated migrationRetirement notice

Cost and performance optimization

DecisionChoiceSignal
Latency too highSmaller model → streaming → shorter max_tokens → caching → fast mode“p95 too slow”, “users wait”
Cost too highSmaller model → caching → Batch → trim output → lower effort“spend is too high”, “reduce cost”
Repeated long prefixPrompt caching, stable-first layout“same system prompt every call”
Bulk, latency-tolerantBatch API (50%)“overnight”, “millions of items”
Mixed difficulty streamCascade cheap→expensive with external validator“most tickets are simple, some hard”
Cutting round trips in multi-tool logicProgrammatic tool calling“many sequential tool calls”
Reducing tokens in long agent runsContext editing (clear tool results)“context keeps growing”, “verbose tool output”
Effort tuninglow/medium mechanical, high default, xhigh hardest“subagent renames files” → low

Reliability and failure handling

DecisionChoiceAnti-pattern rejected
Loop exitBranch on stop_reasonProse parse; iteration cap as primary stop
Truncated output (max_tokens)Continue / raise capTreat as complete
refusal stopExplicit fallback pathGeneric error; blind retry
pause_turnRe-send to continueTreat as failure
Transient 429/5xx/529Backoff + jitter + retry-afterRetry storm; retry 4xx
Tool failureStructured is_error (category, retryable, partial)Generic message; silent empty success
Side-effecting retryIdempotency keyDouble-charge on retry
Model fallbackNewer-or-equal onlyDown-migration drops Fable 5.1 thinking blocks
Escalation triggerExplicit request / beyond capabilitySentiment; self-reported confidence

Data and knowledge management

DecisionChoiceSignal
Small stable corpus fits windowLong context (no retrieval)“50-page handbook, rarely changes”
Large / changing corpus, citationsRAG“thousands of docs”, “must cite”
Style/format not factsFine-tuning / examples“match our tone”, “always this format”
State across sessionsMemory tool / external store“remember across chats”, “survive compaction”
Structured docs, precise + broadParent-child chunking“tables and sections”, “need context around hits”
Vague / multi-intent queryQuery rewriting / multi-query / HyDE“users ask fuzzy questions”
Mixed exact-ID + paraphrase queriesHybrid (dense + BM25) + reranker“order numbers and natural language”
Confident-but-wrong after refreshSuspect retrieval/indexing first“worked before the doc update”

Prompting depth

DecisionChoiceSignal
Format-sensitive / edge casesFew-shot with representative examples“output varies”, “specific format”
Hard reasoning, clean outputAdaptive/extended thinking“show your work but keep answer clean”
Untrusted content in promptXML boundaries, treat as data“user-supplied”, “web content”
Steer output startPrefill (or structured output for guarantees)“always begins with”, “JSON only”
Machine-consumed outputoutput_config.format + validate/retry“downstream system parses it”
Protect main contextSubagent with explicit context“parallel exploration”, “don’t pollute context”

Signal word → principle quick index

Fast lookup: the exact phrase item writers use, and the principle it points to. When two phrases appear in one stem, the constraint (cost, latency, compliance, reliability) usually decides.

Stem phrasePoints to
“MOST cost-effective”Cheapest model / Batch / caching that still meets constraints
“FIRST step”Cheapest reversible diagnostic before big changes
“BEST approach”Simplest design meeting every stated constraint
“Select TWO”Two independently-correct actions; watch for one-right-one-plausible
“at scale” / “millions”Haiku 4.5 + Batch + caching
“sub-second” / “real time”Smaller model, streaming, low effort; not Batch
“overnight” / “latency does not matter”Batch API (50%)
“same long prompt every request”Prompt caching, stable-prefix first
“must never” / “under no circumstances”Deterministic hook / tool permission, not a prompt sentence
“how does the loop know it is done”stop_reason: end_turn, not prose parsing
“when should it escalate”Explicit request or beyond-capability; not sentiment/self-report
“the agent kept going forever”Terminate on stop_reason, not iteration cap
“returns nothing / empty”Structured is_error / not_found; silent-success anti-pattern
“confidence” / “how sure are you”Do not route on self-reported confidence
“angry” / “frustrated customer”Sentiment ≠ complexity; do not escalate on tone
“overall accuracy is 95%”Demand per-segment metrics
“are you sure?” in the same chatFresh session / different model for review
“18 tools” / “too many tools”4–8 tools; tool search + defer_loading
“forced to call a tool” + Fable 5.1auto+instruction / strict / structured output (forced = 400)
“guaranteed JSON shape”output_config.format / strict: true + validation
“thinking blocks disappeared after fallback”Fable 5.1 binding; only migrate up
“context window is full”Context editing (tool results) or compaction (narrative)
“edited an earlier turn” + Fable 5.1Append-only history; use mid-conversation system messages
“remember across sessions”Memory tool / external store
“must cite sources”RAG + citations; provenance test
“worked before the document update”Suspect retrieval/indexing first
“exact IDs and paraphrases”Hybrid retrieval + reranker
“vague user questions”Query rewriting / HyDE / multi-query
“50-page doc that rarely changes”Long context, not RAG
“match our writing style”Few-shot / fine-tuning, not RAG
“reused across several apps/agents”MCP server, not a custom tool
“remote, multi-user tool server”Streamable HTTP + OAuth 2.1, per-user identity
“acts as a shared admin account”Authz gap; propagate end-user identity
“delete/refund tool it doesn’t need”Remove it (least privilege), not log/confirm
“web content / tool output told it to…”Indirect injection; treat as data, boundaries
“hard rule in the system prompt”Prompt-as-enforcement anti-pattern; use a hook
“multi-file / architectural change” (Claude Code)Plan mode
“single obvious edit”Direct mode
“reusable multi-step procedure”Skill
“repo conventions / build commands”CLAUDE.md
“parameterised one-liner”Slash command
“isolated delegated task”Subagent
“dozens–hundreds of agents”Dynamic workflows
“a few delegated tasks”Subagents / forks
“named roles collaborating”Agent teams
“scheduled recurring job”Routines
“distribute commands/agents/hooks/MCP”Plugins
“run in CI / headless”claude -p --output-format json, restricted tools
“grade subjective quality”Rubric → LLM-as-judge (separate session), calibrated
“compare two prompts/models”Pairwise A/B with statistical significance
“gate the deploy”CI regression suite on the golden set
“before and after release”Offline golden set + online canary
“regulated / irreversible / external action”Human-in-the-loop gate
“EU personal data”GDPR; DPIA for high-risk
“health data / PHI”HIPAA + BAA
“US federal”FedRAMP High via Bedrock/Vertex
“prompts must not be retained”ZDR (note Fable 5.1 exclusion)
“non-technical sponsor / executives”ADR + decision matrix + cost model + risk register
“expect 100% accuracy”Set SLO/SLA against measured baselines
“handoff to the team”Runbooks, monitoring, exit criteria
“which model should we use”Match capability matrix to the binding constraint
“reduce hallucinations”Grounding (RAG/citations), validation, not just “tell it not to”

Exam signal

When a stem stacks a capability need against a compliance or cost constraint, the constraint wins. “We want the MCP connector but we are ZDR-only on Fable 5.1” is impossible as stated — recognise the conflict rather than picking the shiny feature.

Last updated Sep 18, 2026