AI Cert Prep
Type to search documentation.

Appendix · Claude

Claude API Cheat Sheet

Messages API request and response anatomy, streaming events, tool use, structured outputs, thinking, caching, batches, errors and retries – with Python and TypeScript.

Request anatomy

json
{
"model": "claude-opus-5",
"max_tokens": 2048,
"system": "You are a precise assistant. Answer only from <document>.",
"messages": [
{ "role": "user", "content": [
{ "type": "text", "text": "<document>…</document>", "cache_control": { "type": "ephemeral" } },
{ "type": "text", "text": "Summarise the termination clause." }
]}
],
"temperature": 0.2,
"stop_sequences": ["</answer>"],
"thinking": { "type": "adaptive" },
"effort": "high",
"tools": [],
"tool_choice": { "type": "auto" },
"metadata": { "user_id": "hashed-id" }
}
FieldNotes
modelPinned ID: claude-fable-5-1, claude-opus-5, claude-sonnet-5, claude-haiku-4-5
max_tokensHard output cap; stop_reason: max_tokens = truncated
systemTop-level string or content blocks; stable → cacheable
messagesAlternating user/assistant; content is string or block array (text, image, document, tool_use, tool_result, thinking)
temperature / top_pSet one, not both; 0 reduces variance, not error
thinking{"type":"adaptive"} (current models); {"type":"enabled","budget_tokens":N} only Haiku 4.5
effortlow / medium / high / xhigh
tools / tool_choiceSee tool use below; forced choice is 400 on Fable 5.1
output_config.formatJSON Schema for structured output

Response anatomy

json
{
"id": "msg_01…",
"type": "message",
"role": "assistant",
"model": "claude-opus-5",
"content": [
{ "type": "thinking", "thinking": "…", "signature": "…" },
{ "type": "text", "text": "The notice period is 60 days." }
],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 1200,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 18000,
"output_tokens": 42
}
}

stop_reason – branch on it every time

ValueMeaningAction
end_turnFinished naturallyDone
tool_useWants tools runExecute, append tool_result, call again
max_tokensTruncatedContinue or raise cap; never treat as complete
stop_sequenceHit a stop stringDone (check stop_sequence)
pause_turnLong server-side tool turn pausedRe-send to continue
refusalSafety declineExplicit fallback path; do not retry blindly

Minimal calls

python
from anthropic import Anthropic
client = Anthropic() # reads ANTHROPIC_API_KEY
msg = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
system="You are a concise analyst.",
messages=[{"role": "user", "content": "Three risks of vendor lock-in?"}],
)
print(msg.content[0].text, msg.stop_reason, msg.usage)

Streaming (SSE)

Event order: message_start → (content_block_start → content_block_delta* → content_block_stop)* → message_delta (carries stop_reason, output usage) → message_stop.

python
with client.messages.stream(
model="claude-sonnet-5", max_tokens=1024,
messages=[{"role": "user", "content": "Write a haiku about latency."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
final = stream.get_final_message()
print(final.stop_reason, final.usage)

Tool use round-trip

json
// 1. Request with tools
{ "model": "claude-opus-5", "max_tokens": 1024,
"tools": [{
"name": "get_order",
"description": "Look up one order by ID. Use when the user references an order number. Returns status, items and total. Does not modify anything.",
"input_schema": { "type": "object",
"properties": { "order_id": { "type": "string", "description": "Order ID, e.g. ORD-12345" } },
"required": ["order_id"] },
"strict": true
}],
"messages": [{ "role": "user", "content": "Where is ORD-12345?" }] }
// 2. Response: stop_reason = "tool_use"
{ "content": [{ "type": "tool_use", "id": "toolu_01", "name": "get_order", "input": { "order_id": "ORD-12345" } }],
"stop_reason": "tool_use" }
// 3. Follow-up with tool_result (in a USER message)
{ "messages": [
{ "role": "user", "content": "Where is ORD-12345?" },
{ "role": "assistant", "content": [{ "type": "tool_use", "id": "toolu_01", "name": "get_order", "input": { "order_id": "ORD-12345" } }] },
{ "role": "user", "content": [{ "type": "tool_result", "tool_use_id": "toolu_01",
"content": "{\"status\":\"shipped\",\"eta\":\"2026-09-17\"}" }] }
] }
// Error result – structured, never empty-success
{ "type": "tool_result", "tool_use_id": "toolu_01", "is_error": true,
"content": "{\"category\":\"not_found\",\"retryable\":false,\"message\":\"No order ORD-12345\"}" }

The agentic loop (Python)

python
def run(messages, tools, model="claude-opus-5"):
while True:
r = client.messages.create(model=model, max_tokens=4096, tools=tools, messages=messages)
messages.append({"role": "assistant", "content": r.content})
if r.stop_reason == "tool_use":
results = []
for block in r.content:
if block.type == "tool_use":
try:
out = TOOLS[block.name](**block.input)
results.append({"type": "tool_result", "tool_use_id": block.id, "content": json.dumps(out)})
except ToolError as e:
results.append({"type": "tool_result", "tool_use_id": block.id, "is_error": True,
"content": json.dumps({"category": e.category, "retryable": e.retryable, "message": str(e)})})
messages.append({"role": "user", "content": results})
continue
if r.stop_reason == "max_tokens":
messages.append({"role": "user", "content": "Continue."}); continue
if r.stop_reason == "pause_turn":
continue
if r.stop_reason == "refusal":
return handle_refusal(r)
return r # end_turn / stop_sequence

tool_choice

ValueBehaviourFable 5.1
{"type":"auto"}Model decides (default)✓
{"type":"any"}Must call some tool400
{"type":"tool","name":"x"}Must call tool x400
{"type":"none"}No tools this turn✓
disable_parallel_tool_use: trueOne tool per turn✓

Structured outputs

json
{ "model": "claude-sonnet-5", "max_tokens": 1024,
"output_config": { "format": { "type": "json_schema", "schema": {
"type": "object",
"properties": {
"vendor": { "type": "string" },
"total": { "type": "number" },
"currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] },
"line_items": { "type": "array", "items": { "type": "object",
"properties": { "sku": { "type": "string" }, "qty": { "type": "integer" } },
"required": ["sku", "qty"], "additionalProperties": false } }
},
"required": ["vendor", "total", "currency", "line_items"],
"additionalProperties": false } } },
"messages": [{ "role": "user", "content": [
{ "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "…" } },
{ "type": "text", "text": "Extract the invoice." } ] }] }

Always validate downstream and run a validation-retry loop that feeds the specific error back.

Prompt caching

json
{ "system": [{ "type": "text", "text": "<20k-token style guide>", "cache_control": { "type": "ephemeral" } }],
"tools": [ … ],
"messages": [ … dynamic content last … ] }
  • Order: tools → system → messages; cache breakpoints mark the end of a stable prefix.
  • Minimum ~1024 tokens (2048 on Haiku 4.5). Up to 4 breakpoints.
  • 5-minute TTL default (write 1.25×); 1-hour option (write 2×). Reads 0.1×.
  • usage.cache_read_input_tokens confirms hits.

Message Batches

python
batch = client.messages.batches.create(requests=[
{"custom_id": f"doc-{i}", "params": {"model": "claude-haiku-4-5", "max_tokens": 512,
"messages": [{"role": "user", "content": doc}]}} for i, doc in enumerate(docs)])
# poll batch.processing_status until "ended", then stream results
for res in client.messages.batches.results(batch.id):
if res.result.type == "succeeded": ...

50% discount; results within 24 h; per-item success/error; ideal for “overnight, cost matters”.

Errors and retries

HTTPTypeRetry?
400invalid_request_errorNo – fix the request (e.g., forced tool_choice on Fable 5.1, budget_tokens on Opus 5)
401authentication_errorNo – key
403permission_errorNo – entitlement
404not_found_errorNo – model/resource
413request_too_largeNo – shrink
429rate_limit_errorYes – backoff, honour retry-after
500api_errorYes – backoff
529overloaded_errorYes – backoff, consider fallback model

Exponential backoff with jitter; idempotency keys for side-effecting tools; log request_id from response headers.

Other inputs and features

FeatureShape
VisionContent block type: image with a source of type base64 or url
PDFContent block type: document with a base64 or url source and media_type: application/pdf
Files APIUpload once, reference by file_id in a content block
CitationsEnable on documents to get source spans
Server-side toolsweb_search, code_execution, text_editor, bash, memory, computer
Tool searchtool_search tool plus defer_loading: true on catalogue tools
MCP connectorTop-level mcp_servers array of objects with type: url, url and name
Context editingcontext_management strategies that clear old tool results server-side
CompactionServer-side summarisation preserving narrative
Memory toolPersistent file-like store across sessions
json
// Vision and PDF content blocks
{ "type": "image", "source": { "type": "url", "url": "https://example.com/chart.png" } }
{ "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "…" } }
// MCP connector
{ "mcp_servers": [{ "type": "url", "url": "https://mcp.example.com/mcp", "name": "orders" }] }

Full streaming event sequence

The wire format is SSE. A complete turn with one text block and one tool call looks like this (elided deltas marked …):

text
event: message_start
data: {"type":"message_start","message":{"id":"msg_01","role":"assistant","model":"claude-opus-5","content":[],"stop_reason":null,"usage":{"input_tokens":1200,"output_tokens":1}}}
event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"thinking","thinking":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"thinking_delta","thinking":"Checking the order…"}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"signature_delta","signature":"Er8B…"}}
event: content_block_stop
data: {"type":"content_block_stop","index":0}
event: content_block_start
data: {"type":"content_block_start","index":1,"content_block":{"type":"text","text":""}}
event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"text_delta","text":"Looking that up"}}
event: content_block_stop
data: {"type":"content_block_stop","index":1}
event: content_block_start
data: {"type":"content_block_start","index":2,"content_block":{"type":"tool_use","id":"toolu_01","name":"get_order","input":{}}}
event: content_block_delta
data: {"type":"content_block_delta","index":2,"delta":{"type":"input_json_delta","partial_json":"{\"order_id\":\"ORD-12345\"}"}}
event: content_block_stop
data: {"type":"content_block_stop","index":2}
event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"tool_use","stop_sequence":null},"usage":{"output_tokens":57}}
event: message_stop
data: {"type":"message_stop"}
event: ping (may arrive at any time; ignore)
EventCarriesNotes
message_startShell message, input usagecontent is empty; stop_reason null
content_block_startBlock type at an indexOne per text/thinking/tool_use block
content_block_deltatext_delta, thinking_delta, signature_delta, input_json_deltaTool input streams as partial JSON – buffer per index and parse at stop
content_block_stopBlock finished–
message_deltastop_reason and cumulative output usageBranch here, not on prose
message_stopTurn complete–
pingKeep-aliveIgnore
erroroverloaded_error etc. mid-streamHandle like the HTTP error

Streaming pitfall

Tool input arrives as input_json_delta fragments; never JSON.parse a partial. Accumulate partial_json per block index and parse only after content_block_stop. The authoritative stop_reason is on message_delta.

tool_result with images

A tool can return an image (e.g., a rendered chart) alongside text. The content of a tool_result accepts a block array:

json
{ "role": "user", "content": [
{ "type": "tool_result", "tool_use_id": "toolu_09", "content": [
{ "type": "text", "text": "Chart rendered for Q3 revenue." },
{ "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgo…" } }
] }
]}

Parallel tool calls – full round trip

The model can emit several tool_use blocks in one turn. Execute them concurrently and return all tool_result blocks in the next single user message.

json
// Assistant turn: three parallel calls (stop_reason: tool_use)
{ "role": "assistant", "content": [
{ "type": "tool_use", "id": "toolu_a", "name": "get_weather", "input": { "city": "Paris" } },
{ "type": "tool_use", "id": "toolu_b", "name": "get_weather", "input": { "city": "Tokyo" } },
{ "type": "tool_use", "id": "toolu_c", "name": "get_fx", "input": { "pair": "EURJPY" } }
]}
// Your next user turn: all results, matched by tool_use_id, order-independent
{ "role": "user", "content": [
{ "type": "tool_result", "tool_use_id": "toolu_a", "content": "{\"c\":18}" },
{ "type": "tool_result", "tool_use_id": "toolu_b", "content": "{\"c\":26}" },
{ "type": "tool_result", "tool_use_id": "toolu_c", "content": "{\"rate\":171.2}" }
]}

Use disable_parallel_tool_use: true in tool_choice to force one call per turn when calls have side effects that must be ordered.

Structured output with validation-retry

python
import json, jsonschema
from anthropic import Anthropic
client = Anthropic()
SCHEMA = { "type": "object",
"properties": { "vendor": {"type": "string"}, "total": {"type": "number"},
"currency": {"type": "string", "enum": ["USD","EUR","GBP"]} },
"required": ["vendor","total","currency"], "additionalProperties": False }
def extract(doc_text, max_attempts=3):
messages = [{"role": "user", "content": f"Extract the invoice.\n<doc>{doc_text}</doc>"}]
for attempt in range(max_attempts):
r = client.messages.create(
model="claude-sonnet-5", max_tokens=1024,
output_config={"format": {"type": "json_schema", "schema": SCHEMA}},
messages=messages)
text = r.content[0].text
try:
data = json.loads(text)
jsonschema.validate(data, SCHEMA) # schema + business rules
if data["total"] < 0:
raise ValueError("total must be non-negative")
return data
except (json.JSONDecodeError, jsonschema.ValidationError, ValueError) as e:
messages += [{"role": "assistant", "content": text},
{"role": "user", "content": f"That failed validation: {e}. Return corrected JSON only."}]
raise RuntimeError("extraction failed after retries") # escalate, never silently return bad data

The loop feeds the specific error back and caps attempts, then escalates — never returns an empty or unvalidated object (silent-failure anti-pattern).

Prompt caching – multi-breakpoint layout

Up to four breakpoints. Order most-stable → least-stable so the longest possible prefix stays cached when only the tail changes.

json
{
"system": [
{ "type": "text", "text": "<static company style guide, 15k tokens>", "cache_control": { "type": "ephemeral" } }
],
"tools": [
{ "name": "search_kb", "description": "…", "input_schema": { }, "cache_control": { "type": "ephemeral" } }
],
"messages": [
{ "role": "user", "content": [
{ "type": "text", "text": "<retrieved policy docs, changes per session>", "cache_control": { "type": "ephemeral", "ttl": "1h" } },
{ "type": "text", "text": "Now: the user's current question (never cached)." }
]}
]
}
BreakpointContentTTLRationale
1System style guide5-minNever changes; deepest prefix
2Tool definitions5-minStable across the app
3Session documents1-hourReused all session; longer TTL earns the 2× write
—Current questionnoneUnique per request

Rule: a breakpoint caches everything before it. Putting a volatile block early invalidates all deeper caching.

Batch lifecycle states

text
create → in_progress ──► (per request: succeeded | errored | canceled | expired)
└─► ended (all requests terminal; results retrievable)
cancel ─► canceling ─► ended
processing_statusMeaning
in_progressStill running; poll request_counts
cancelingCancellation requested
endedTerminal; fetch results stream
Per-request result.typeHandling
succeededUse result.message
erroredInspect result.error; may re-submit that item
canceledBatch was canceled before this ran
expiredNot completed within the 24 h window; re-submit

Results are available for 29 days. Match items by custom_id; do not assume order.

Error handling with backoff

python
import time, random
from anthropic import Anthropic, APIStatusError, RateLimitError, APIConnectionError
client = Anthropic()
RETRYABLE = {429, 500, 502, 503, 529}
def call_with_retry(**params):
for attempt in range(6):
try:
return client.messages.create(**params)
except RateLimitError as e:
wait = float(e.response.headers.get("retry-after", 0)) or min(60, 2 ** attempt)
time.sleep(wait + random.uniform(0, 0.5))
except APIStatusError as e:
if e.status_code in RETRYABLE:
time.sleep(min(60, 2 ** attempt) + random.uniform(0, 0.5))
else:
raise # 400/401/403/404/413 → fix, don't retry
except APIConnectionError:
time.sleep(min(60, 2 ** attempt) + random.uniform(0, 0.5))
raise RuntimeError("exhausted retries")

The SDKs retry automatically with backoff; a custom loop matters when you tune the ceiling, add jitter, or switch to a fallback model on repeated 529.

Rate-limit headers and idempotency

HeaderMeaning
anthropic-ratelimit-requests-remainingRPM budget left
anthropic-ratelimit-input-tokens-remainingITPM budget left
anthropic-ratelimit-output-tokens-remainingOTPM budget left
anthropic-ratelimit-*-resetWhen each bucket refills (RFC 3339)
retry-afterSeconds to wait after a 429/529 – honour it
request-idLog this for every request; include in support tickets

Proactively throttle when a *-remaining header approaches zero rather than waiting for the 429.

python
# Idempotency: safe retries for side-effecting requests
client.messages.create(**params, extra_headers={"idempotency-key": f"charge-{order_id}"})

Reuse the same key on retries so a duplicate delivery does not double-charge. Design side-effecting tools the same way (accept an idempotency key argument).

Files API and citations

python
# Upload once, reference by file_id across many requests
f = client.files.upload(file=("contract.pdf", open("contract.pdf","rb"), "application/pdf"))
r = client.messages.create(
model="claude-sonnet-5", max_tokens=1024,
messages=[{"role": "user", "content": [
{"type": "document", "source": {"type": "file", "file_id": f.id},
"citations": {"enabled": True}},
{"type": "text", "text": "What is the termination notice period? Cite the clause."}
]}])

With citations enabled, text blocks carry a citations array of source spans:

json
{ "type": "text", "text": "The notice period is 60 days.",
"citations": [{ "type": "page_location", "cited_text": "…sixty (60) days…",
"document_index": 0, "start_page_number": 4, "end_page_number": 4 }] }

Citations enable the provenance test: every claim points back to a span you can open.

Context editing and compaction request shapes

json
// Context editing: clear stale tool results server-side, keep the turn valid
{ "model": "claude-opus-5", "max_tokens": 4096,
"context_management": {
"edits": [{ "type": "clear_tool_uses", "trigger": { "type": "input_tokens", "value": 100000 },
"keep": { "type": "tool_uses", "value": 3 } }]
},
"messages": [ … ] }
json
// Compaction: summarise older turns while preserving the narrative
{ "context_management": {
"edits": [{ "type": "compact", "trigger": { "type": "input_tokens", "value": 150000 } }] } }
StrategyRemovesKeepsUse when
Context editing (clear_tool_uses)Verbose old tool resultsRecent N tool uses, all textLong tool-heavy agent runs
Compaction (compact)Old turns → summaryNarrative continuityLong conversational sessions

Both run server-side, so they keep Fable 5.1’s append-only history valid (client-side trimming would not).

Memory tool

json
{ "model": "claude-opus-5", "max_tokens": 2048,
"tools": [{ "type": "memory_20250818", "name": "memory" }],
"messages": [{ "role": "user", "content": "Remember I prefer metric units, then convert 5 miles." }] }

The model reads/writes a persistent file-like store (via memory tool calls you execute against your backing store) that survives compaction and new sessions. Use for durable preferences and state; do not stuff it into the system prompt.

MCP connector (server-side)

python
r = client.messages.create(
model="claude-opus-5", max_tokens=1024,
mcp_servers=[{ "type": "url", "url": "https://mcp.example.com/mcp", "name": "orders",
"authorization_token": user_scoped_token }],
extra_headers={"anthropic-beta": "mcp-client-2025-04-04"},
messages=[{"role": "user", "content": "Where is ORD-12345?"}])

Anthropic’s infrastructure connects to the remote MCP server for you — no local client harness. Pass a user-scoped token so tools act as the end user, not a shared super-user.

Managed Agents vs Agent SDK

Managed AgentsAgent SDK (claude-agent-sdk)
Who runs the loopAnthropic (hosted loop + sandbox)You (your infra)
Ops burdenMinimalYou own scaling, sandboxing, secrets
Control over runtime/network/data localityLimitedFull
Tools/MCP/hooks/subagentsConfiguredFull programmatic control
Choose when“least operational overhead”, “don’t want to host”“control the runtime”, “data must stay in our VPC”, “custom harness”
python
# Agent SDK sketch – you host the harness
from claude_agent_sdk import ClaudeAgent
agent = ClaudeAgent(model="claude-opus-5", tools=[...], mcp_servers=[...],
permission_mode="acceptEdits")
result = agent.run("Refactor the auth module and run the tests.")

Common misconceptions

MisconceptionRealityWhy it matters on the exam
“tool_result goes in an assistant message”It goes in a user message, matched by tool_use_idWrong-role distractor
“Parse the text for ‘done’ to end the loop”Branch on stop_reasonProse-parsing anti-pattern
“max_tokens truncation is a finished answer”It is truncation — continue or raise the capSilent-failure distractor
“Retry every error”Only 429/5xx/529 with backoff; fix 4xxRetry-storm distractor
“Structured output means you can skip validation”Still validate + retry business rulesOver-trust distractor
“Caching a volatile block early saves money”It invalidates every deeper cache; stable firstCaching-layout distractor
“Empty result is fine when a tool finds nothing”Return structured is_error/not_foundSilent empty-success anti-pattern

Scenario walkthrough

A team runs an agent that calls 3–4 tools per turn (some parallel), on Opus 5, in a long session that occasionally hits 529s and grows past 150k tokens. Payments tools must not double-charge. What does a correct implementation look like?

  1. Loop control — branch on stop_reason; on tool_use execute all parallel tool_use blocks concurrently and return every tool_result in one user message.
  2. Ordering side effects — for the payment tool, set disable_parallel_tool_use when it must be sequenced, and pass an idempotency key so a retried call is safe.
  3. 529 handling — exponential backoff with jitter honouring retry-after; after repeated 529, fall back to a newer-or-equal model (Opus 5 → Fable 5.1 is up-safe; never down to an older model mid-session if thinking blocks matter).
  4. Context growth — configure clear_tool_uses context editing at ~100k input tokens keeping the last 3 tool uses; server-side so history stays valid.
  5. Errors from tools — structured is_error results with category/retryable, never empty success.
  6. Observability — log request-id and usage per call; watch anthropic-ratelimit-*-remaining and throttle before 429.

Rejected alternatives: parsing prose to stop (prose-parsing), retrying 400s (retry-storm), client-side history trimming (breaks append-only), and self-reported “I finished” as the loop exit (self-report reliance).

Last updated Sep 18, 2026