AI Cert Prep
Type to search documentation.

Domains

D3 · Agents and Workflows

Workflow vs agent decision criteria, the agentic loop, subagents and hierarchies, the Claude Agent SDK, managed vs self-hosted agents, hooks, memory, context management, frameworks, and safe termination.

This domain is roughly 8 of 53 items. It tests whether you can choose between a fixed workflow and an autonomous agent, build the agentic loop correctly (driven by stop_reason, terminated safely), structure subagents with isolated context, and use the Claude Agent SDK and hooks. The recurring theme: structure and determinism where it matters; autonomy only where it pays.

Learning objectives

By the end of this page you should be able to:

  1. Decide between a workflow and an agent, and pick the right Anthropic pattern.
  2. Implement the agentic loop terminated by stop_reason, not iteration caps.
  3. Design manager/supervisor hierarchies and subagents with isolated context.
  4. Use the Claude Agent SDK (Python + TS) and its key options.
  5. Distinguish Managed Agents (hosted) from self-hosted loops/harnesses.
  6. Use hooks for deterministic actions, the memory tool, and manage the context window in long agents.
  7. Position frameworks (LangGraph, PydanticAI, CrewAI) and apply safe termination.

3.1 Workflow vs agent

  • A workflow orchestrates Claude and tools through predefined code paths – predictable, testable, cheap.
  • An agent lets Claude dynamically direct its own process and tool use – flexible, but less predictable and harder to test.

Prefer the simplest thing that works: use workflows for well-defined, repeatable tasks; use agents only when the path cannot be predicted in advance.

SignalChoose
Fixed steps, known inputs/outputsWorkflow
Predictable branchingWorkflow (routing)
Open-ended task, unknown number of stepsAgent
Need auditability and low costWorkflow
Task requires the model to decide the planAgent

3.2 The Anthropic ‘Building effective agents’ patterns

PatternWhat it isUse when
Prompt chainingOutput of step N feeds step N+1Task decomposes into fixed sequential subtasks
RoutingClassify input, dispatch to a specialised pathDistinct categories need different handling
ParallelizationRun subtasks concurrently, aggregateIndependent subtasks or voting
Orchestrator-workersA lead splits work and delegates to workersSubtasks unknown until runtime
Evaluator-optimizerOne generates, another critiques and refinesQuality improves with iteration and clear criteria
Autonomous agentModel plans and acts in a loop with toolsOpen-ended, path not predictable

Exam signal

“Steps are fixed” → chaining. “Different categories” → routing. “Independent subtasks / vote” → parallelization. “Lead delegates dynamically” → orchestrator-workers. “Generate then critique” → evaluator-optimizer. “Open-ended, decide its own steps” → autonomous agent.


3.3 The agentic loop

The loop is: call the model → if stop_reason == 'tool_use', run the tools, append tool_result, call again → repeat until end_turn.

python
def run_agent(client, messages, tools):
while True:
resp = client.messages.create(
model="claude-opus-5", max_tokens=2048, tools=tools, messages=messages)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason == "tool_use":
tool_results = []
for block in resp.content:
if block.type == "tool_use":
out = dispatch(block.name, block.input) # run the tool
tool_results.append({"type": "tool_result",
"tool_use_id": block.id, "content": out})
messages.append({"role": "user", "content": tool_results})
continue
return resp # end_turn / refusal / max_tokens handled by caller
text
┌─ model call ─┐
│ │ stop_reason == tool_use
│ Claude ────┼──────────────► run tool(s) ──► append tool_result ──┐
│ │ │
└──────────────┘◄────────────────────────────────────────────────── ┘
│ stop_reason == end_turn
▼
return

Safe termination

Terminate on stop_reason, not on an arbitrary iteration count. An iteration cap is a safety backstop against runaway loops, never the primary stopping mechanism (anti-patterns #1 and #2). Also handle max_tokens, pause_turn and refusal.


3.4 Subagents and manager/supervisor hierarchies

A manager (orchestrator) agent decomposes a task and delegates to subagents, each with its own isolated context window, tools and system prompt. Benefits:

  • Context isolation – a subagent’s noisy intermediate work does not pollute the manager’s context.
  • Specialisation – each subagent has a focused tool allowlist and prompt.
  • Parallelism – independent subagents run concurrently.
text
┌──────────── Manager / Coordinator ────────────┐
│ plans, delegates, aggregates results │
└───┬───────────────┬───────────────┬───────────┘
▼ ▼ ▼
Subagent A Subagent B Subagent C
(own context) (own context) (own context)

Exam signal

“Intermediate research is bloating the context”, “specialised subtasks”, “run parts in parallel” → subagents with isolated context under a coordinator.


3.5 The Claude Agent SDK

pip install claude-agent-sdk / npm install @anthropic-ai/claude-agent-sdk (renamed from the Claude Code SDK). It provides the agent loop, tool handling and Claude Code’s harness.

python
import anyio
from claude_agent_sdk import query, ClaudeAgentOptions
async def main():
options = ClaudeAgentOptions(
system_prompt="You are a careful coding assistant.",
allowed_tools=["Read", "Grep", "Edit"],
permission_mode="acceptEdits",
mcp_servers={"docs": {"command": "python", "args": ["docs_server.py"]}},
max_turns=8,
)
async for message in query(prompt="Fix the failing test in test_utils.py", options=options):
print(message)
anyio.run(main)

Key options: allowed_tools (allowlist – least privilege), permission_mode (default/acceptEdits/plan/bypassPermissions), system_prompt, mcp_servers, hooks (deterministic handlers), max_turns (backstop, not primary termination).


3.6 Managed vs self-hosted agents

Managed AgentsAgent SDK / Tool Runner (self-hosted)
Who runs the loopAnthropic hosts loop + sandboxYou host
Control over environmentLowerFull
Ops burdenMinimalYou manage infra, sandboxing, scaling
Best forFast start, standard agentic tasksCustom tools, custom environments, data control

Exam signal

“We want Anthropic to run the sandbox and loop” → Managed Agents. “We need our own tools/environment/data controls” → self-hosted via the Agent SDK.


3.7 Hooks, memory and context management

  • Hooks are deterministic shell/HTTP handlers fired at lifecycle points (PreToolUse, PostToolUse, UserPromptSubmit, Stop, SessionStart, Notification, SubagentStop, PreCompact). Exit code 2 blocks the action. Use them to enforce rules that must not depend on the model’s cooperation.
  • Memory tool persists information across sessions (durable notes, learned facts) rather than re-deriving it each run.
  • Context-window management in long agents: prune old tool results (context editing), summarise while preserving narrative (compaction), and offload durable state to memory. Without this, long agents blow the window and cost.
python
# A PreToolUse hook that blocks destructive shell commands (exit 2 = block)
# hook script: reads JSON on stdin, exits 2 to deny
import json, sys
event = json.load(sys.stdin)
cmd = event.get("tool_input", {}).get("command", "")
if "rm -rf" in cmd:
print("Blocked: destructive command", file=sys.stderr)
sys.exit(2)
sys.exit(0)

Deterministic enforcement

Critical rules (never delete, never spend, always approve) belong in hooks, not the prompt (anti-pattern #3). A hook cannot be talked out of blocking; a prompt instruction can.


3.8 Frameworks (positioning)

FrameworkPosition
Claude Agent SDKAnthropic’s first-party harness; tightest Claude Code / tool integration
LangGraphGraph-based orchestration of nodes/edges; explicit state machines
PydanticAIType-safe, Pydantic-validated agent outputs in Python
CrewAIMulti-agent “crews” with roles and tasks

Frameworks add orchestration and ergonomics; they do not change the fundamentals – you still drive the loop from stop_reason, enforce critical rules deterministically, and manage context.


3.9 A complete agentic loop dispatching on every stop_reason

The sketch in 3.3 only handles tool_use and end_turn. A production loop must dispatch on every stop_reason: tool_use, end_turn, max_tokens, pause_turn, refusal (and stop_sequence). The iteration cap is a backstop, never the primary stop.

python
from anthropic import Anthropic
client = Anthropic()
def run_agent(messages, tools, max_iters=25):
for _ in range(max_iters): # backstop only, not the primary stop
resp = client.messages.create(
model="claude-opus-5", max_tokens=4096, tools=tools, messages=messages)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason == "tool_use":
results = []
for block in resp.content:
if block.type == "tool_use":
out = dispatch(block.name, block.input) # run the tool
results.append({"type": "tool_result",
"tool_use_id": block.id, "content": out})
messages.append({"role": "user", "content": results})
continue
if resp.stop_reason == "pause_turn":
# long-running server tool paused; resend to resume
continue
if resp.stop_reason == "max_tokens":
# truncated output — ask to continue or raise max_tokens; do not treat as done
messages.append({"role": "user", "content": "Continue the previous response."})
continue
if resp.stop_reason == "refusal":
# model declined for safety; surface to a human, do not retry blindly
return {"status": "refused", "response": resp}
if resp.stop_reason in ("end_turn", "stop_sequence"):
return {"status": "done", "response": resp}
raise RuntimeError("iteration backstop hit — investigate; loop did not reach end_turn")

Dispatch on the signal, not the prose

Treating max_tokens as completion silently truncates work; treating refusal as an error to retry blindly ignores a safety signal; ignoring pause_turn abandons a long-running server tool. Parsing the assistant text for ‘I am done’ is anti-pattern #1. Always branch on stop_reason.


3.10 Orchestrator-workers with subagent context passing

In an orchestrator-workers pattern the lead agent decomposes the task, spawns workers each with an isolated context window, passes each worker only the slice it needs, and aggregates the distilled results back — the workers’ verbose intermediate work never enters the orchestrator’s context.

python
from anthropic import Anthropic
client = Anthropic()
def worker(task_brief: str, documents: str) -> str:
"""A subagent with its OWN fresh context — only the brief + its slice, nothing else."""
resp = client.messages.create(
model="claude-sonnet-5", max_tokens=1024,
system="You are a focused research worker. Return only a distilled 5-line summary.",
messages=[{"role": "user", "content": f"<task>{task_brief}</task>\n<docs>{documents}</docs>"}])
return "".join(b.text for b in resp.content if b.type == "text")
def orchestrator(question: str, corpus: dict[str, str]) -> str:
# 1. Lead decomposes and delegates; each worker gets only its slice (context passing).
summaries = []
for topic, docs in corpus.items():
brief = f"Summarise what the docs say about: {question} (focus: {topic})"
summaries.append(f"[{topic}] {worker(brief, docs)}") # only distilled result returns
# 2. Lead aggregates distilled summaries — worker noise never entered this context.
synthesis = client.messages.create(
model="claude-opus-5", max_tokens=2048,
system="You are the coordinator. Synthesise the worker summaries into one answer.",
messages=[{"role": "user", "content": f"Question: {question}\n\n" + "\n".join(summaries)}])
return "".join(b.text for b in synthesis.content if b.type == "text")
text
┌──────── Orchestrator (Opus 5) ────────┐
│ decompose → delegate slices → synthesise│
└───┬───────────┬───────────┬────────────┘
brief+slice brief+slice brief+slice (context passing: only the needed slice)
▼ ▼ ▼
Worker A Worker B Worker C (each fresh, isolated context)
│ │ │
5-line 5-line 5-line (only distilled results return)
summary summary summary

Exam signal

‘A lead splits work it cannot enumerate up front and delegates’ → orchestrator-workers. ‘Pass each subagent only the slice it needs and return distilled summaries’ → context passing + isolation, which keeps the coordinator’s window clean and cheap.


3.11 Hooks in the Agent SDK: a PreToolUse example

Hooks are deterministic handlers fired at lifecycle points; exit code 2 blocks the action. Critical rules belong here, not in the prompt (anti-pattern #3). Below, a PreToolUse hook denies any Bash command containing a destructive pattern before it ever runs.

python
import anyio
from claude_agent_sdk import query, ClaudeAgentOptions
async def block_destructive(input_data, tool_use_id, context):
"""PreToolUse hook: return a deny decision to block the tool call."""
cmd = input_data.get("tool_input", {}).get("command", "")
if any(p in cmd for p in ("rm -rf", "DROP TABLE", "git push --force")):
return {
"hookSpecificOutput": {
"hookEventName": "PreToolUse",
"permissionDecision": "deny",
"permissionDecisionReason": "Blocked destructive command by policy.",
}
}
return {}
async def main():
options = ClaudeAgentOptions(
allowed_tools=["Read", "Grep", "Bash"],
permission_mode="default",
hooks={"PreToolUse": [block_destructive]},
max_turns=8,
)
async for message in query(prompt="Clean up the build artifacts", options=options):
print(message)
anyio.run(main)

Deterministic enforcement, not prompt pleading

A hook cannot be argued out of blocking; a system-prompt rule can be overridden by a clever tool result or injected instruction. Any rule that is irreversible or high-stakes (delete, spend, publish, push) must be a hook or programmatic check (anti-pattern #3).


3.12 Managed Agents vs Agent SDK vs custom loop

DimensionManaged AgentsClaude Agent SDK (self-hosted)Custom Messages API loop
Who runs the loop + sandboxAnthropic hosts bothYou host; SDK provides the harnessYou build and host everything
Setup effortLowestModerateHighest
Control over environment / toolsLowestHigh (own tools, MCP, hooks, permission modes)Total
Built-in hooks, permissions, subagentsYes (managed)Yes (SDK primitives)You implement it all
Data / infra controlAnthropic-managedYour infra, your data controlsYour infra, your data controls
Best forFast start on standard agentic tasksCustom tools/environments needing Claude Code harness ergonomicsFull control, unusual integrations, or minimal dependencies
Termination / safety burdenManagedYou wire hooks + stop_reason + max_turnsYou implement stop_reason dispatch and backstops yourself

Exam signal

‘Anthropic should host the loop and sandbox, start fast’ → Managed Agents. ‘We need our own tools/environment/data controls with a first-party harness, hooks and permission modes’ → Agent SDK. ‘Minimal dependencies / unusual integration / full control’ → custom Messages API loop (and you must implement stop_reason dispatch and a backstop yourself).


3.13 The evaluator-optimizer pattern in code

Evaluator-optimizer improves quality by iterating: a generator produces a draft, a separate evaluator critiques it against explicit criteria, and the generator revises — repeating until the criteria are met or a backstop trips. The evaluator must not share the generator’s session (anti-pattern #9).

python
from anthropic import Anthropic
client = Anthropic()
def generate(brief, feedback=None):
prompt = brief if not feedback else f"{brief}\n\nRevise to fix: {feedback}"
r = client.messages.create(model="claude-sonnet-5", max_tokens=1024,
messages=[{"role": "user", "content": prompt}])
return "".join(b.text for b in r.content if b.type == "text")
def evaluate(brief, draft):
# Separate session AND ideally a different model — no shared reasoning bias.
r = client.messages.create(
model="claude-opus-5", max_tokens=512,
system="Score the draft against the brief. Return JSON {'pass': bool, 'feedback': str}.",
messages=[{"role": "user", "content": f"BRIEF:\n{brief}\n\nDRAFT:\n{draft}"}])
import json
return json.loads("".join(b.text for b in r.content if b.type == "text"))
def optimize(brief, max_rounds=3): # backstop, not the primary stop
draft = generate(brief)
for _ in range(max_rounds):
verdict = evaluate(brief, draft)
if verdict["pass"]:
return draft # primary stop: criteria met
draft = generate(brief, verdict["feedback"])
return draft # backstop reached — flag for review
ElementRequirement
Generator and evaluatorDifferent sessions; ideally different models
StoppingPrimary: criteria met; backstop: max rounds
FeedbackConcrete and actionable, fed back into the next draft
Use whenQuality improves with iteration and criteria are explicit

Exam signal

‘Generate, then critique and refine against clear criteria’ → evaluator-optimizer. If the same session grades its own draft, that is anti-pattern #9 — the judge inherits the generator’s bias.


3.14 Positioning frameworks against the fundamentals

Frameworks add orchestration ergonomics; they never remove the need for stop_reason-driven termination, deterministic enforcement, and context management.

FrameworkSweet spotWhat it does not change
Claude Agent SDKFirst-party harness; tightest Claude Code / tool / hook integrationYou still wire allowlists, hooks, stop_reason dispatch, backstop
LangGraphExplicit graph/state-machine orchestration of nodes and edgesStill drive the model call from stop_reason inside each node
PydanticAIType-safe, Pydantic-validated agent outputs in PythonValidation-retry and schema design still apply
CrewAIMulti-agent ‘crews’ with roles and tasksLeast privilege, context isolation, and safe termination still apply

A framework is not a safety control

Choosing LangGraph/CrewAI/PydanticAI does not enforce critical rules or terminate loops for you. Enforcement is still hooks/permissions; termination is still stop_reason with a backstop. Any answer implying ‘use framework X so we do not need to handle termination/enforcement’ is wrong.


3.15 Common misconceptions

MisconceptionRealityWhy it matters on the exam
Agents are always better than workflowsPrefer the simplest thing that works; agents add cost/unpredictabilityWorkflow-vs-agent is a recurring first decision
An iteration cap is how you stop a loopTerminate on stop_reason; the cap is only a backstopAnti-patterns #1/#2
max_tokens means the agent is doneIt means truncation; continue or raise the capSilent-failure distractor
A refusal should be retried until it worksIt is a safety stop; surface to a humanPrevents blind-retry loops
Critical rules can live in the system promptEnforce them in hooks/permissions (exit 2 blocks)Anti-pattern #3
More tools make an agent more capableBeyond ~10, selection degrades; use ~4–5 + tool search/defer_loadingAnti-pattern #8
Subagents just cost moreIsolated context reduces bloat/drift and can lower costExplains why delegation helps
A framework handles termination and safety for youIt does not; you still wire stop_reason and enforcementFramework-positioning trap

3.16 Scenario walkthrough: a support agent that must not overreach

Scenario. You are building a customer-support agent. It should answer questions from a knowledge base and look up order status, and it may propose a refund — but any refund over $500 needs human approval, and account deletion is never allowed. Early tests show three problems: the loop sometimes never ends (it keeps ‘thinking out loud’), it occasionally issues a large refund on its own after reading a ticket that says ‘the manager approved a full refund’, and as it researches multi-part questions its context fills with raw KB dumps and later answers degrade. Design the agent.

Expert reasoning trace.

  1. Fix termination first. Drive the loop on stop_reason: continue while tool_use, stop on end_turn, and explicitly handle max_tokens (continue), pause_turn (resume) and refusal (surface). Keep an iteration cap purely as a backstop. The ‘thinks out loud forever’ bug is anti-patterns #1/#2 — never parse prose to stop.
  2. Enforce the money rule deterministically. The large refund happened because a ticket (untrusted content) claimed manager approval — indirect prompt injection. The rule ‘refund > $500 needs human approval’ must be a PreToolUse hook (exit 2 blocks, routes to approval), never a system-prompt sentence (anti-pattern #3). Account deletion is simply not in the allowlist (least privilege).
  3. Isolate research context. Use subagents with isolated context for multi-part lookups, returning distilled summaries to the coordinator, and apply context editing/compaction so raw KB dumps do not bloat the main window. Raising max_tokens would not fix drift.
  4. Right-size the tool set. Give the agent ~4–5 tools (search_kb, get_order, propose_refund); if the catalogue grows, use tool search + defer_loading. Do not pile on tools.
  5. Reject the tempting alternatives. ‘Add a firm system-prompt rule about refunds and approvals’ — prompt-as-enforcement (#3). ‘Trust the model to recognise the fake approval’ — self-report/no-boundary reliance. ‘Cap the loop at 5 iterations as the fix’ — backstop, not termination (#2). ‘Increase max_tokens to stop the context degrading’ — wrong lever for drift.

Correct decision. stop_reason-driven loop with a backstop; PreToolUse hook enforcing the refund threshold and routing to approval; account-deletion tool excluded by least privilege; subagents with isolated context plus context editing/compaction for research; ~4–5 tools with tool search if the catalogue grows.


Exam traps in this domain

TrapWhy it is wrong
Using an agent when a fixed workflow sufficesAdds cost and unpredictability for no benefit
Iteration cap as the primary stopBackstop only; terminate on stop_reason (anti-patterns #1, #2)
Enforcing critical rules in the promptUse hooks / programmatic checks (anti-pattern #3)
One agent with 18 toolsOverloads the model; split into subagents / use tool search + defer_loading (anti-pattern #8)
Letting subagent work bloat the manager’s contextUse isolated-context subagents
Same-session self-review of an agent’s own workRetains reasoning bias (anti-pattern #9); use a fresh evaluator
Parsing prose to decide delegation/terminationDrive from structured signals
Assuming a framework removes the need for safe terminationIt does not
Treating max_tokens as task completionOutput was truncated; ask to continue or raise the cap, do not mark done
Retrying a refusal blindlyIt is a safety signal; surface to a human, do not loop on it
Ignoring pause_turnAbandons a long-running server tool; resend to resume
Passing full worker context back to the orchestratorBloats the coordinator window; return only distilled results
Having the same session grade the agent’s own draft in evaluator-optimizerJudge inherits the generator’s bias (anti-pattern #9); use a separate session/model
Assuming a framework (LangGraph/CrewAI/PydanticAI) handles termination or enforcementIt does not; you still wire stop_reason dispatch and hooks/permissions
Acting on an ‘approval’ claimed inside untrusted ticket/tool contentIndirect injection; enforce approval in a hook, not by trusting the content
Choosing an agent when fixed ordered steps existPrompt chaining is cheaper, testable and predictable
Using an iteration cap as the fix for a runaway loopThe cap is a backstop; the fix is stop_reason-driven termination

Practice questions

Q1 · A task has three fixed, ordered steps with known inputs and outputs. What is the best design? (Select one)

A. An autonomous agent that decides the steps. B. A prompt-chaining workflow where each step feeds the next. C. A single mega-prompt. D. An orchestrator with dynamic subagents.

Answer: B. Fixed sequential subtasks are the definition of prompt chaining – predictable, testable, cheap. Autonomy (A, D) adds unpredictability with no benefit; a mega-prompt (C) is harder to control.

Q2 · An agent loop sometimes never terminates. What is the correct primary termination mechanism? (Select one)

A. Stop after 5 iterations regardless. B. Stop when the assistant text says it is finished. C. Continue while stop_reason == 'tool_use' and stop on end_turn, with an iteration cap only as a safety backstop. D. Stop when output length exceeds a threshold.

Answer: C. Terminate on stop_reason; the cap is a backstop, not the primary mechanism (anti-patterns #1, #2). Prose (B) and length (D) are not reliable signals.

Q3 · A research agent's context is filling with verbose intermediate tool output, degrading later steps. What is the best fix? (Select one)

A. Increase max_tokens. B. Delegate research to subagents with isolated context that return only distilled results to the coordinator. C. Lower temperature. D. Remove all tools.

Answer: B. Isolated-context subagents keep noisy intermediate work out of the coordinator’s window and return summaries. max_tokens (A) caps output; temperature (C) and removing tools (D) do not address context bloat.

Q4 · A business rule states the agent must never execute a refund over $500 without human approval. Where should this be enforced? (Select one)

A. In the system prompt as an instruction. B. In a PreToolUse hook that blocks (exit code 2) refunds over the threshold and routes to human approval. C. By asking the model to double-check. D. By lowering the model’s temperature.

Answer: B. Critical, irreversible rules must be enforced deterministically via hooks, not prompt instructions (anti-pattern #3). Prompt-based enforcement and self-checks can be bypassed.

Q5 · A team wants Anthropic to host the agent loop and sandbox so they can start fast with standard tools. Which option fits? (Select one)

A. Self-hosted Agent SDK on their own infrastructure. B. Managed Agents. C. A raw Messages API loop with no tools. D. A cron job.

Answer: B. Managed Agents have Anthropic host the loop and sandbox. Self-hosting (A) is the opposite; a bare loop (C) or cron (D) does not provide a managed sandbox.

Q6 · Which TWO Agent SDK options directly support least-privilege and safety? (Select two)

A. allowed_tools restricting the tool allowlist. B. max_turns set to 1000. C. hooks that block dangerous actions. D. system_prompt length. E. permission_mode: 'bypassPermissions'.

Answer: A and C. An explicit tool allowlist and blocking hooks enforce least privilege and safety. A huge max_turns (B) weakens the backstop; prompt length (D) is irrelevant; bypassPermissions (E) removes safeguards.

Q7 · One agent is configured with 18 tools and frequently picks the wrong one. What is the recommended remedy? (Select one)

A. Add more tools to cover edge cases. B. Reduce to a focused set (≈4–5), split responsibilities into subagents, and use tool search with defer_loading for large catalogues. C. Raise temperature. D. Increase max_turns.

Answer: B. Too many tools per agent is anti-pattern #8; the fix is fewer tools, subagent specialisation, and tool search + defer_loading beyond ~10 tools. More tools (A) worsens it.

Q8 · An agent must remember user preferences across separate sessions. Which mechanism is designed for this? (Select one)

A. Increasing the context window. B. The memory tool for cross-session persistence. C. Higher effort thinking. D. Re-sending the full transcript every time forever.

Answer: B. The memory tool persists durable information across sessions. A bigger window (A) does not persist between sessions; effort (C) is unrelated; re-sending everything (D) is costly and unbounded.

Q9 · An agent loop returns `stop_reason == 'max_tokens'` mid-way through a long answer, and the harness treats that as completion. What is the correct handling? (Select one)

A. Treat max_tokens as done; the answer is complete. B. Recognise the output was truncated: ask the model to continue (or raise max_tokens) rather than marking the turn complete. C. Retry the whole request at temperature: 0. D. Switch to Haiku 4.5.

Answer: B. max_tokens means the output hit the output cap and was cut off — it is not end_turn. The loop should continue the response or raise the cap. Treating it as done (A) silently truncates work; regenerating (C) or switching model (D) does not recover the truncated content.

Q10 · While running, an agent receives `stop_reason == 'pause_turn'`. What does this indicate and what should the loop do? (Select one)

A. The model refused; stop and alert a human. B. A long-running server-side tool paused; resend the conversation to resume it. C. The iteration cap was hit; abort. D. The output was truncated; raise max_tokens.

Answer: B. pause_turn signals a long-running (server) turn paused; the loop resends to resume. Refusal (A) is a different stop reason; the cap (C) is a harness backstop, not a stop_reason; truncation (D) is max_tokens.

Q11 · A research coordinator delegates topics to worker subagents. Which TWO design choices keep the coordinator's context clean and costs down? (Select two)

A. Give each worker its own isolated context window with only its task slice. B. Return each worker’s full transcript, including intermediate tool output, to the coordinator. C. Have workers return only a short distilled summary to the coordinator. D. Run all work in the coordinator’s single context. E. Give every worker all 18 tools.

Answer: A and C. Isolated worker contexts plus distilled-summary returns keep the coordinator’s window free of noisy intermediate work and reduce token cost. Returning full transcripts (B) or using one shared context (D) causes bloat; giving every worker 18 tools (E) is anti-pattern #8.

Q12 · A team must guarantee the agent never runs `git push --force`, regardless of what the model decides. Where and how should this be enforced in the Agent SDK? (Select one)

A. A strongly worded instruction in the system prompt. B. A PreToolUse hook that returns a deny decision (or a shell hook exiting with code 2) when the command matches the pattern. C. Ask the model to confirm before force-pushing. D. Lower the temperature so it behaves.

Answer: B. Irreversible, critical rules must be enforced deterministically with a PreToolUse hook (deny / exit code 2), which the model cannot argue around (anti-pattern #3). Prompt instructions (A), model self-confirmation (C), and temperature (D) can all be bypassed.

Q13 · An organisation wants the tightest control over its own tools, sandboxing and data, using a first-party harness with hooks and permission modes. Which option fits, and what stays their responsibility? (Select one)

A. Managed Agents; Anthropic handles termination and safety entirely. B. The Claude Agent SDK self-hosted; they still wire allowed_tools, hooks, stop_reason dispatch and a max_turns backstop themselves. C. A cron job invoking a single prompt. D. A bare Messages API call with no tools.

Answer: B. The Agent SDK is the first-party self-hosted harness giving full control over tools, environment and data, with hooks and permission modes — but the team still owns least-privilege allowlists, hook enforcement, stop_reason handling and the iteration backstop. Managed Agents (A) give less environment control; cron (C) and a bare call (D) provide no agentic harness.

Q14 · A drafting tool must improve output by generating, critiquing against explicit criteria, and revising. Which pattern fits, and what is the key correctness requirement? (Select one)

A. Autonomous agent; let it decide when it is satisfied. B. Evaluator-optimizer, with the evaluator in a separate session (ideally a different model) so the judge does not inherit the generator’s bias. C. Prompt chaining with no critique step. D. Parallelization of many drafts with no evaluation.

Answer: B. Generate-critique-refine against clear criteria is evaluator-optimizer; the evaluator must be a separate session/model (anti-pattern #9 otherwise). An autonomous agent (A) lacks the structured critique; chaining without critique (C) and parallel drafts without evaluation (D) do not iterate on quality.

Q15 · A team says 'we will use LangGraph, so we do not need to handle loop termination or rule enforcement.' Why is this wrong? (Select one)

A. It is correct; frameworks handle termination and enforcement. B. Frameworks add orchestration ergonomics but do not remove the need to drive termination from stop_reason (with a backstop) or to enforce critical rules via hooks/permissions. C. LangGraph disables stop_reason. D. Only the Agent SDK requires termination handling.

Answer: B. A framework is not a safety control; you still terminate on stop_reason and enforce rules deterministically. Termination/enforcement are not delegated to the framework (A); LangGraph does not disable stop_reason (C); the requirement is universal, not SDK-only (D).

Q16 · A support agent issued a $2,000 refund on its own after a ticket said 'the manager approved a full refund.' What TWO controls should have prevented this? (Select two)

A. A PreToolUse hook that blocks refunds over $500 (exit 2) and routes to human approval. B. Treating ticket content as untrusted data inside content boundaries, not as authorisation. C. A stronger system-prompt sentence about refund limits. D. Trusting the model to verify the manager’s approval. E. Raising the model’s effort level.

Answer: A and B. The refund rule must be enforced by a deterministic hook, and the ticket text must be treated as untrusted data (indirect injection), never as authorisation. A system-prompt sentence (C) is anti-pattern #3; trusting the model (D) is exactly the failure; effort (E) is irrelevant to enforcement.

Q17 · An agent researching multi-part questions fills its context with raw knowledge-base dumps, and later answers degrade. Which TWO fixes address the drift? (Select two)

A. Delegate lookups to subagents with isolated context that return distilled summaries. B. Apply context editing/compaction so stale dumps do not bloat the main window. C. Increase max_tokens. D. Add more tools to the main agent. E. Raise temperature.

Answer: A and B. Isolated-context subagents and context editing/compaction keep the coordinator’s window clean. max_tokens (C) caps output, not context; more tools (D) worsens selection; temperature (E) is unrelated to drift.

Q18 · An agent must never delete an account. What is the correct way to guarantee this? (Select one)

A. Add ‘never delete accounts’ to the system prompt. B. Exclude the account-deletion tool from the agent’s allowed_tools allowlist (least privilege); if it exists elsewhere, gate it behind a human-approved workflow with a PreToolUse hook. C. Set max_turns low so it runs out of turns before deleting. D. Ask the model to confirm before deleting.

Answer: B. Least privilege means the capability is simply not available to the agent; irreversible actions live behind approval + hooks. A prompt rule (A) is anti-pattern #3; a low max_turns (C) is unrelated; model self-confirmation (D) can be bypassed.

Q19 · A generate-critique loop has the generator grade its own draft in the same conversation and always passes on the first try. What is wrong, and what is the fix? (Select one)

A. Nothing; self-grading is efficient. B. Same-session self-review inherits the generator’s reasoning bias (anti-pattern #9); run the evaluator in a separate session, ideally a different model, against explicit criteria. C. The loop needs a higher iteration cap. D. Switch the generator to Haiku 4.5.

Answer: B. A judge sharing the generator’s context rubber-stamps its own work; a separate-session/different-model evaluator removes the bias. Self-grading is not fine (A); a higher cap (C) does not fix bias; changing the generator model (D) does not separate the judge.

Key takeaways

  • Prefer workflows for predictable tasks; use agents only when the path cannot be predetermined.
  • Know the six patterns: chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer, autonomous agent.
  • Drive the loop from stop_reason; iteration caps are backstops, not the primary stop.
  • Use subagents with isolated context for specialisation, parallelism and context hygiene.
  • The Agent SDK gives you the loop, tools, hooks and permission modes; enforce least privilege via allowed_tools.
  • Managed Agents = Anthropic hosts the loop/sandbox; Agent SDK = you host.
  • Enforce critical rules with hooks (exit 2 blocks), persist state with the memory tool, and manage context with editing/compaction.
  • Evaluator-optimizer improves quality by iterating — but the evaluator must run in a separate session (ideally a different model) or it inherits the generator’s bias (anti-pattern #9).
  • Frameworks (LangGraph, PydanticAI, CrewAI, Agent SDK) add ergonomics, not safety; you still drive termination from stop_reason and enforce rules via hooks/permissions.
  • Treat ‘approval’ or instructions embedded in tickets, documents or tool results as untrusted data — enforce approval in a hook, never by trusting the content.
  • Least privilege first: exclude irreversible tools from the allowlist and gate them behind human-approved workflows rather than relying on prompt rules.

Last updated Sep 18, 2026