The 10 Anti-Patterns
The ten critical anti-patterns that appear as wrong answers on the Architect exam — what each looks like in a stem, why it fails, the correct alternative with a snippet, and how it is worded as a distractor.
Anthropic publishes ten anti-patterns that recur as wrong answers on the Architect and Developer exams. If you can spot them in an option, you can eliminate that option instantly. For each: what it looks like in a question stem, why it fails, the correct alternative with a snippet, and the phrasing examiners use to make it sound right.
The distractors are engineered to be tempting — they usually appear “reasonable”, “simpler”, or “more flexible”. The table below is your fast reference; the sections expand each one.
| # | Anti-pattern | Correct alternative |
|---|---|---|
| 1 | Parsing natural language for loop termination | Check stop_reason |
| 2 | Arbitrary iteration caps as the primary stop | stop_reason primary; cap as backstop |
| 3 | Prompt-based enforcement of critical rules | Deterministic hooks (exit code 2) |
| 4 | Self-reported confidence for escalation | Explicit request or capability gap |
| 5 | Sentiment-based escalation | Sentiment ≠ complexity |
| 6 | Generic error messages hiding context | Structured errors (category, retryable) |
| 7 | Silently suppressing errors as success | Surface failures explicitly |
| 8 | Too many tools per agent | 4–5 focused; tool search + defer_loading |
| 9 | Same-session self-review | Independent evaluator (fresh session/model) |
| 10 | Aggregate metrics masking per-type failure | Per-segment metrics, gate on worst |
Anti-pattern 1 · Parsing natural language for loop termination
In a stem: “The agent loop stops when Claude’s response contains the phrase ‘task complete’.” / “We detect completion by scanning the output text for ‘done’.”
Why it fails: The model’s prose is not a control signal. It varies with phrasing, can be wrong, and can be manipulated by prompt injection in tool results. Loops that key off text stop early, never stop, or stop when hijacked.
Correct alternative: Drive the loop from the API’s stop_reason.
if resp.stop_reason == "tool_use": ... # run tools, continueelif resp.stop_reason == "end_turn": break # doneelif resp.stop_reason == "max_tokens": raise OutputTruncated() # not completionDistractor wording: “For flexibility, terminate when the model says it has finished” — sounds natural-language-friendly and adaptive.
Anti-pattern 2 · Arbitrary iteration caps as the primary stop
In a stem: “The loop runs a fixed 10 iterations and then stops.” / “We limit the agent to 5 tool calls to keep it bounded.”
Why it fails: A cap alone tells you nothing about whether the task finished. If the loop only ends by hitting the cap, results are silently incomplete; if the cap is too high, cost runs away.
Correct alternative: stop_reason is the primary stop; the cap is a backstop against runaway loops.
iterations = 0while resp.stop_reason == "tool_use" and iterations < MAX_ITERS: # cap = safety net ... iterations += 1Distractor wording: “To prevent infinite loops, cap iterations and treat reaching the cap as completion” — conflates a safety net with a completion signal.
Anti-pattern 3 · Prompt-based enforcement of critical rules
In a stem: “We added ‘never run destructive commands’ to the system prompt.” / “CLAUDE.md tells Claude to always run tests before committing.”
Why it fails: Prompt instructions are probabilistic and vulnerable to injection. A critical business or safety rule that must always hold cannot rely on best-effort compliance.
Correct alternative: Enforce with a deterministic hook (exit code 2 blocks) or programmatic guard.
# PreToolUse hook: block the action deterministicallyif echo "$cmd" | grep -q 'rm -rf'; then echo "blocked" >&2; exit 2; fiDistractor wording: “Strengthen the instruction and use an emphatic, all-caps rule in the prompt” — implies wording can make a probabilistic control deterministic.
Anti-pattern 4 · Self-reported confidence for escalation
In a stem: “The agent escalates when its self-reported confidence drops below 70%.”
Why it fails: Language models are poorly calibrated; their stated confidence does not reliably track correctness. Routing on it produces both false escalations and missed ones.
Correct alternative: Escalate on explicit user request (immediately) or an objective capability gap (after attempting resolution).
if turn.user_explicitly_requested_human: escalate() # nowelif turn.needs_capability_agent_lacks and turn.attempts_exhausted: escalate()Distractor wording: “Use the model’s confidence score to decide when a human is needed” — sounds data-driven and principled.
Anti-pattern 5 · Sentiment-based escalation
In a stem: “When sentiment analysis flags the customer as frustrated, escalate to a human.”
Why it fails: Sentiment is not complexity. A frustrated customer may have a trivial, resolvable request; a calm one may have a genuinely hard problem. Sentiment-driven routing sends resolvable cases to humans and can miss complex ones.
Correct alternative: Handle emotion with a good response; escalate on explicit request or capability gap. Use sentiment to adjust tone, not to route.
Distractor wording: “Improve customer experience by escalating angry customers immediately” — appeals to empathy.
Anti-pattern 6 · Generic error messages hiding context
In a stem: “On failure the tool returns ‘An error occurred’.” / “Errors are caught and a generic message is shown.”
Why it fails: The agent (and operators) lose the diagnostic detail needed to decide whether to retry, escalate, or fix the input. Recovery becomes guesswork.
Correct alternative: Return structured errors with category, retryable flag, message and any partial results.
{ "status": "error", "category": "rate_limit", "retryable": true, "retry_after": 12, "partial": [] }Distractor wording: “For a clean user experience, hide technical details behind a friendly generic message” — sounds like good UX.
Anti-pattern 7 · Silently suppressing errors as success
In a stem: “On failure the tool returns an empty list.” / “Exceptions are swallowed and the agent continues.”
Why it fails: A failure disguised as “no results” or success propagates wrong data downstream. The agent cannot distinguish a real empty result from a broken call.
Correct alternative: Surface failures explicitly; distinguish {status:"ok", items:[]} from {status:"error", …}.
try: return {"status": "ok", "items": query()}except BackendError as e: return {"status": "error", "category": "backend", "retryable": True, "message": str(e)}Distractor wording: “To keep the agent robust, catch errors and return empty results so the flow never breaks” — sounds resilient.
Anti-pattern 8 · Too many tools per agent
In a stem: “The agent is configured with 18 tools covering every operation.”
Why it fails: Each tool schema consumes context, and past a handful the model’s selection accuracy drops — it picks the wrong tool or hallucinates arguments.
Correct alternative: Give each agent 4–5 focused tools; split responsibilities across subagents; for large catalogues enable the tool search tool and defer_loading: true.
tools = [ {"type": "tool_search_tool_20250101", "name": "tool_search"}, {"name": "search_kb", "defer_loading": True, "input_schema": {...}}, # loads on demand]Distractor wording: “Give the agent all available tools so it never lacks a capability” — sounds thorough and flexible.
Anti-pattern 9 · Same-session self-review
In a stem: “We ask the model, in the same conversation, to check whether its answer is correct.”
Why it fails: The reviewing turn shares the reasoning context that produced the error, so it inherits the same bias and tends to confirm the original answer.
Correct alternative: Use an independent evaluator — a fresh session, ideally a different model — scoring against an explicit rubric (evaluator-optimizer, LLM-as-judge).
judge = client.messages.create(model="claude-sonnet-5", # different model/session messages=[{"role": "user", "content": f"<rubric>{rubric}</rubric><answer>{ans}</answer> Score it."}])Distractor wording: “Have the model self-critique in the same chat to catch its own mistakes” — sounds efficient.
Anti-pattern 10 · Aggregate metrics masking per-type failure
In a stem: “Extraction accuracy is 94% overall, so the pipeline is ready.” (while one document type is at 60%)
Why it fails: A single aggregate number hides a category that is failing badly. The worst-performing segment is exactly where the business risk lives.
Correct alternative: Report per-segment / per-document-type metrics and gate on the worst type, not the average.
overall: 94% ← misleadinginvoices: 99% contracts: 60% receipts: 96% ← the real picture; gate on contractsDistractor wording: “The aggregate accuracy exceeds our 90% bar, so ship it” — sounds metric-driven and objective.
Distractor families: how examiners disguise each anti-pattern
The ten anti-patterns cluster into families that share a “sound good” surface. Recognising the family lets you eliminate two or three options at once.
| Family | Anti-patterns | Surface appeal | The tell |
|---|---|---|---|
| Control-flow-by-vibes | #1 prose parsing, #2 iteration cap | “flexible”, “prevents infinite loops” | Correctness must come from stop_reason |
| Enforcement-by-hope | #3 prompt enforcement | “just add a clear instruction” | Critical rules need a deterministic hook |
| Trust-the-model’s-feelings | #4 confidence, #5 sentiment | “data-driven”, “empathetic” | Escalate on explicit request or capability gap |
| Hide-the-failure | #6 generic errors, #7 silent success | “clean UX”, “robust” | Surface structured errors; distinguish empty from failure |
| More-is-better | #8 too many tools | “thorough”, “never lacks a capability” | 4–5 focused tools; search + defer beyond ~10 |
| Grade-your-own-homework | #9 same-session review | “efficient self-check” | Independent evaluator, fresh session/model |
| One-number-to-rule-them | #10 aggregate metrics | “objective”, “meets the bar” | Slice per segment; gate on the worst |
Two-anti-pattern stems
Harder items place two anti-patterns as separate options (e.g. one distractor escalates on sentiment (#5) and another on confidence (#4)). Eliminating one is not enough — recognise the family and reject both, then pick the explicit-request/capability-gap answer.
Worked distractor elimination
Stem: “A support agent should hand off to a human. The customer is angry but their request (order status) is fully resolvable, and the model reports 62% confidence. Which is correct?”
A. Escalate due to negative sentiment. → #5 sentiment ✗ eliminateB. Escalate because confidence < 70%. → #4 self-report ✗ eliminateC. Resolve the request; neither sentiment nor → explicit-request/capability rule ✓ confidence is a valid trigger, and it is within capability.D. Cap the chat at 3 turns then escalate. → arbitrary cap, not an escalation rule ✗Two distractors (A, B) belong to the trust-the-model’s-feelings family; D is a control-flow-by-vibes cap misapplied to escalation. C applies the actual rule.
How to use this list in the exam
Q1 · An option reads: 'Cap the agent at 10 iterations and treat hitting the cap as task completion.' Which anti-pattern is this? (Select one)
A. Anti-pattern 1. B. Anti-pattern 2 — iteration cap as the primary/completion signal. C. Anti-pattern 8. D. Anti-pattern 10.
Answer: B. Treating the cap as completion is anti-pattern 2. The cap is only a backstop; completion comes from stop_reason.
Q2 · An option reads: 'Escalate whenever the sentiment classifier flags frustration.' Which two anti-patterns are related distractor families here, and which one is this exactly? (Select one)
A. This is anti-pattern 5 (sentiment); its cousin is anti-pattern 4 (self-reported confidence). B. This is anti-pattern 4; its cousin is anti-pattern 6. C. This is anti-pattern 9. D. This is anti-pattern 7.
Answer: A. Sentiment-based escalation is #5; the related escalation trap is #4 (self-reported confidence). Correct escalation is explicit request or capability gap.
Q3 · An option reads: 'On any backend error, return an empty result so the pipeline keeps running.' Which anti-pattern, and what is the fix? (Select one)
A. Anti-pattern 6; add a generic message. B. Anti-pattern 7; return a structured error distinguishing empty-success from failure so downstream code and the agent can react. C. Anti-pattern 3; use a hook. D. Anti-pattern 10; add per-type metrics.
Answer: B. Returning empty on failure is silent suppression (#7). The fix is structured errors that never disguise a failure as ‘no results’.
Q4 · A stem offers: 'Give the agent all 18 available tools so it never lacks a capability.' Which anti-pattern, and what is the correct alternative? (Select one)
A. #10; report per-segment metrics.
B. #8 (too many tools per agent); give 4–5 focused tools, split responsibilities to subagents, or enable tool search with defer_loading beyond ~10.
C. #1; check stop_reason.
D. #3; use a hook.
Answer: B. ‘All tools for completeness’ is #8 — it bloats context and degrades selection. The fix is fewer focused tools, subagent split, or tool search + defer_loading. The other options name unrelated anti-patterns.
Q5 · An option reads: 'Have the model, in the same chat, grade whether its answer meets the rubric.' Which anti-pattern, and the fix? (Select one)
A. #9 (same-session self-review); run an independent evaluator (fresh session, ideally a different model) against the rubric. B. #4; use explicit-request escalation. C. #2; add an iteration cap. D. #6; return a structured error.
Answer: A. In-session self-grading inherits the generator’s bias (#9); independence removes it. The other anti-patterns are unrelated to evaluation bias.
Q6 · A single stem offers two escalation distractors — one on sentiment, one on self-reported confidence — plus a resolvable request. Which TWO options must you eliminate, and why? (Select two)
A. Escalate on negative sentiment. B. Resolve the request, since it is within capability and neither trigger applies. C. Escalate on confidence below a threshold. D. Escalate on explicit request for a human. E. Escalate because the reply is long.
Answer: A and C. Sentiment (#5) and self-reported confidence (#4) are both invalid triggers and belong to the same distractor family — eliminate both. Resolving (B) is correct here; explicit request (D) would be valid but is not present; length (E) is irrelevant.
Q7 · An option reads: 'Aggregate accuracy is 94% and exceeds our 90% bar, so ship it.' What is wrong, and what should gate the release? (Select one)
A. Nothing; 94% exceeds the bar. B. #10 — the aggregate can hide a segment (e.g. contracts at 60%); report per-document-type accuracy and gate on the worst-performing type. C. #7 — return empty on failure. D. #2 — add an iteration cap.
Answer: B. A single aggregate number masks a failing segment (#10); gate on the worst type, not the average. C and D name unrelated anti-patterns, and shipping on the aggregate (A) is the trap itself.
Last updated Sep 18, 2026