AI Cert Prep
Type to search documentation.

Domains

D5 · Context Management and Reliability

Context window economics, context editing vs compaction, memory and external state, subagent isolation, prompt caching, Fable 5.1 append-only history, retries and backoff, graceful degradation and fallback models, provenance, escalation gates, monitoring, rate limits and timeouts.

This domain is worth roughly 9 of 60 items and underpins every scenario: a system that works in a demo but degrades as the context window fills, or falls over when the API returns 429, is not production-ready. It tests the economics of the context window, the two ways to reclaim it (context editing vs compaction), what must live in durable state, and the reliability patterns – caching, retries, fallback, monitoring – that keep an agentic system running.

Learning objectives

By the end of this page you should be able to:

  1. Reason about context window economics and what consumes context.
  2. Choose between context editing (clear tool results) and compaction (summarise preserving narrative).
  3. Use the memory tool and external state for anything that must survive compaction, and subagent isolation to protect the main context.
  4. Use prompt caching as a reliability and cost lever, and honour Fable 5.1’s append-only constraint.
  5. Design retries/backoff/idempotency, graceful degradation and fallback models – and know what breaks when falling back from Fable 5.1.
  6. Maintain provenance and citations, and design escalation / human-in-the-loop gates.
  7. Set up monitoring (latency p95, error rate, cache-hit rate, cost per task) and design for rate limits, timeouts and partial results.

5.1 Context window economics

ModelContext windowMax output
Claude Fable 5.11M128k
Claude Opus 51M128k
Claude Sonnet 51M128k
Claude Haiku 4.5200k64k

The window is a budget, not free space. Everything shares it:

text
┌───────── context window budget ─────────┐
│ system prompt │ stable
│ tool definitions │ stable
│ conversation history (turns) │ grows every turn
│ tool results (can be huge) │ grows fast
│ thinking tokens │ grows with reasoning
└──────────────────────────────────────────┘

As the window fills, latency and cost rise and quality can degrade (“context rot”). The architect’s job is to keep the working set small: cache the stable parts, evict what is no longer needed, and push durable facts to external state.

Exam signal

“Long-running agent whose quality degrades over time” or “context is filling up” points to context editing / compaction / subagent isolation / memory, not a bigger model. Haiku 4.5’s window is 200k, not 1M – a common distractor.


5.2 Context editing vs compaction

Two distinct mechanisms reclaim context. Choosing correctly is tested.

MechanismWhat it doesPreservesUse when
Context editingClears old tool results (and similar bulky blocks) from the windowThe conversation narrative; drops stale tool outputTool results are large and no longer needed, but the dialogue matters
CompactionSummarises the conversation server-side, preserving the narrative in condensed formThe gist/narrative of the whole conversationThe conversation itself has grown too long
text
Context editing: [sys][tools][turn1][BIG tool result ✗ cleared][turn2]… ← keeps dialogue, drops bulk
Compaction: [sys][tools][==== summary of turns 1..40 ====][turn41]… ← condenses the narrative

Exam signal

“Bulky tool outputs we no longer need” → context editing (clear tool results). “The whole conversation is too long but we must keep its thread” → compaction (summarise preserving narrative). Neither preserves everything – anything that must survive verbatim goes to external state or the memory tool.


5.3 Memory tool and external state

Compaction and context editing are lossy. Anything that must survive – a decision, a customer fact, a checkpoint, canonical business data – belongs in durable state:

  • Memory tool – model-managed persistence across sessions.
  • Files – artefacts and checkpoints you control.
  • Database – canonical, transactional business state (orders, tickets).

The conversation references durable state; it does not own it. This is the same rule as Domain 1’s state-vs-conversation table.


5.4 Subagent isolation protects the main context

A subagent runs in its own context window. Delegating a bulky sub-task (reading 50 files, searching a large corpus) to a subagent means the raw material never enters the coordinator’s window – only the distilled result returns. This is one of the most effective context-protection patterns, and it is why explicit context passing (Domain 1) matters: the subagent starts clean.


5.5 Long-context ordering and prompt caching

  • Ordering. Put stable, reusable content first (system, tools, long reference docs); put the variable task last. This both enables caching and keeps the model’s attention on the right material.
  • Prompt caching as a reliability + cost lever. Cache reads ≈ 0.1× base input price; writes ≈ 1.25× (5-min TTL) or 2× (1-hour TTL). Mark the last stable block with cache_control: {"type": "ephemeral"}. Minimum cacheable prefix ~1024 tokens (2048 on Haiku). A high cache-hit rate lowers cost and latency, improving reliability under load.
python
system = [
{"type": "text", "text": SYSTEM_RULES},
{"type": "text", "text": TOOLS_DOC},
{"type": "text", "text": REFERENCE, "cache_control": {"type": "ephemeral"}}, # cache boundary
]

5.6 Fable 5.1: append-only history and thinking-block binding

Fable 5.1 introduces breaking changes an architect must design around:

  • Append-only history. Editing, reordering or removing earlier turns invalidates later thinking blocks. Harnesses must be append-only: freeze system and tools, put mid-session changes in role: "system" messages, and trim server-side via context editing/compaction rather than rewriting the transcript.
  • Thinking-block binding. Thinking blocks are readable only by the producing model or a newer one. Falling back to an older model silently drops those thinking blocks.
  • Operational constraints. Requires 30-day data retention; not in the Priority Tier; thinking always on.

Design implication

On Fable 5.1 you cannot “clean up” the transcript by editing old turns. Design the harness to only append, and reclaim space with context editing/compaction – which operate server-side without invalidating the append-only chain.


5.7 Retries, backoff, idempotency

Retry the transient errors with exponential backoff and jitter, honouring retry-after:

StatusMeaningRetry?
400 invalid_requestBad requestNo — fix the request
401 / 403Auth / permissionNo
413 request_too_largeToo bigNo — reduce size
429 rate_limitRate limitedYes — backoff, respect retry-after
500 api_error / 529 overloadedServer-sideYes — backoff + jitter

Only retry idempotent operations safely; give writes an idempotency key (Domain 1).

python
def call_with_backoff(fn, max_attempts=5):
for i in range(max_attempts):
try:
return fn()
except (RateLimitError, APIError, OverloadedError) as e:
if i == max_attempts - 1:
raise
delay = min(2 ** i, 30) + random.random() # backoff + jitter
time.sleep(getattr(e, "retry_after", delay))

5.8 Graceful degradation and fallback models

When the primary model is unavailable or overloaded, degrade gracefully: fall back to another model, serve a cached/partial result, or queue for later. Fallback has costs.

FallbackConsideration
Fable 5.1 → older modelThinking blocks are silently dropped; behaviour changes; forced-tool-choice rules differ; validate the fallback path explicitly
Opus 5 → Sonnet 5Cheaper/faster, some capability loss; Sonnet 5 disallows mid-conversation system messages and task budgets
Any → Haiku 4.5200k window (not 1M); still uses budget_tokens; big context inputs may not fit

Falling back from Fable 5.1

Because thinking blocks are readable only by the producing or a newer model, a fallback to an older model drops them and can change behaviour subtly. Test the degraded path; do not assume the fallback is a drop-in replacement.


5.9 Provenance and citations

For any output that will be relied on, preserve where it came from: keep citations from web search / retrieval, attach source IDs to extracted facts, and surface provenance so a human can verify. Provenance is both a quality and a governance control (and pairs with Domain 3’s per-type metrics and validation).


5.10 Escalation and human-in-the-loop gates

Irreversible or high-stakes actions need a human gate – the same escalation logic as Domain 1 (explicit request → immediate; capability gap → after attempting; never sentiment/self-report), plus mandatory human approval before irreversible effects (payments, deletions, external sends). The gate is deterministic (a permission/hook), not a prompt request.


5.11 Monitoring

Instrument the system with the metrics that predict failure:

MetricWhy it matters
Latency p50 / p95p95 catches the tail users actually feel
Error rate by categoryDistinguishes rate-limit from validation from model failures
Cache-hit rateLow hit rate silently inflates cost and latency
Cost per taskThe unit economics that decide viability
Retry counts / timeoutsRising retries signal upstream trouble

Pair with per-agent traces and correlation IDs (Domain 1).


5.12 Rate limits, timeouts and partial results

  • Rate limits are per-model-tier (RPM, ITPM, OTPM). Design concurrency and batching to stay under them; use the Batch API (50% discount, within 24h) for latency-tolerant work.
  • Timeouts. Set them at every boundary (tool calls, subagents, the overall task); a hung tool must not hang the agent.
  • Partial results. On timeout or partial failure, return what you have with provenance of the gap (Domain 1’s partial-failure handling), rather than nothing or a false-complete answer.

5.13 Context economics with token arithmetic

Treat the window as a budget you can compute against. A worked example for a long-running agent on Opus 5 (1M window, $5/M in, $25/M out):

text
Turn budget as the session grows (input tokens carried each turn):
system + tools (stable) = 12,000
CLAUDE-style reference docs = 8,000
conversation history @ turn 30 = 60,000
accumulated tool results = 400,000 ← the runaway
thinking (this turn) = 10,000
----------------------------------------------
input this turn = 490,000 tokens
cost of THIS input turn = 490,000 × $5 / 1e6 = $2.45 (and rising every turn)

The accumulated tool results dominate. Three levers, with their arithmetic:

  • Context editing clears the 400k of stale tool results → input drops to ~90k → this-turn input cost falls from $2.45 to ~$0.45.
  • Prompt caching the stable 20k prefix (system + tools + docs) → those tokens read at ~0.1× (20,000 × $0.5/1e6 = $0.01 vs $0.10).
  • Subagent isolation: delegate the tool-heavy sub-task so the 400k never enters the main window at all — only the distilled result returns.
SymptomMechanismRough effect
Bulky, stale tool resultsContext editingRemoves the largest term
Whole dialogue too longCompactionCondenses history to a summary
Repeated stable prefixPrompt caching~10× cheaper reads on the prefix
Sub-task needs huge raw inputSubagent isolationRaw material never hits main window

Exam signal

“Context is filling / cost rising every turn” is almost never solved by “a bigger model”. The math points to context editing (remove the biggest term), caching (cheapen the stable term), or subagent isolation (keep the big term out). Haiku 4.5’s window is 200k, so it is smaller, not a fix.


5.14 The Fable 5.1 harness: constraints and a compliant design

Fable 5.1 (claude-fable-5-1, 1M/128k, $10/$50) has the strictest harness rules on the exam. Design around all of them:

ConstraintConsequenceCompliant design
Append-only historyEditing/reordering/removing earlier turns invalidates later thinking blocksNever rewrite the transcript; only append
Thinking-block bindingThinking readable only by the producing model or newerA fallback to an older model silently drops thinking → test the degraded path
No forced tool_choiceany / {type:"tool"} return 400Use auto + instruction, strict tools, or structured outputs
Thinking always onCannot disable thinkingBudget for thinking tokens; no budget_tokens (Haiku-only)
30-day retention, no ZDRNot zero-data-retentionExclude for ZDR-required workloads
Not Priority TierNo priority capacity guaranteesPlan capacity/fallback accordingly
text
Compliant Fable 5.1 harness:
┌ freeze system + tools (never edited) ─────────────┐
│ append user/assistant turns only │
│ mid-session change? → append a role:"system" msg │ (NOT edit an old turn)
│ window pressure? → server-side context editing │
│ / compaction (append-safe) │
│ need structured out?→ output_config.format / strict │ (never forced tool_choice)
└────────────────────────────────────────────────────┘

Do not “tidy” a Fable 5.1 transcript

Rewriting old turns to remove noise breaks thinking-block binding and produces inconsistent later responses. Reclaim space with server-side context editing/compaction, which do not invalidate the append-only chain.


5.15 Reliability budgets: timeouts, concurrency and rate-limit math

Reliability is quantitative. A quick capacity check for a real-time workload:

text
Tier limits (example): RPM = 4,000 ITPM = 400,000 input tokens/min
Per request: ~10,000 input tokens
Token-bound throughput: 400,000 / 10,000 = 40 requests/min ← ITPM binds first
Request-bound: 4,000 RPM
=> Effective ceiling = min(40, 4,000) = 40 rpm; the token limit dominates.

If you need 10,000 latency-tolerant jobs, the real-time ceiling (40/min ≈ 4+ hours and rate-limit risk) argues for the Batch API (50% off, ≤24h) instead. Set timeouts at every boundary (tool, subagent, overall task) so one hung call cannot stall the system, and on timeout return partial results with provenance of the gap.

BoundaryTimeoutOn breach
Single tool callsecondsStructured timeout error, retryable
Subagenttens of secondsPartial result + gap provenance
Overall tasktask-appropriateEscalate or serve partial with a note

Common misconceptions

MisconceptionRealityWhy it matters on the exam
“A bigger model fixes a filling context.”The fix is context editing/compaction/subagents/caching; a bigger model just delays it.The top D5 distractor.
“Haiku 4.5 has a 1M window.”Haiku 4.5 is 200k; large inputs may not fit.Fallback/window items.
“Compaction and context editing are the same.”Compaction summarises the narrative; context editing clears tool results.Mechanism-choice items.
“Must-survive facts are safe in the conversation.”Compaction/editing are lossy; use the memory tool or a database.State-vs-conversation items.
“Fallback from Fable 5.1 is a drop-in.”Thinking blocks are dropped on older fallbacks; behaviour changes.Graceful-degradation items.
“Retry any error with backoff.”Only 429/5xx/529 are retryable; 400/401/403/413 must be fixed.Retry-classification items.
“Average latency is enough to monitor.”p95/p99 catch the tail users feel; averages hide it.Monitoring items.
“Editing old turns tidies a Fable 5.1 session.”It breaks thinking-block binding; the harness must be append-only.Fable 5.1 harness items.

Scenario walkthrough — a research agent that is fast in a demo and falls over in production

Situation. A multi-agent research agent works in demos but in production (1) degrades after ~30 turns as the window fills with large search results it no longer needs, while the dialogue thread must stay intact; (2) during an Anthropic capacity event it fails over from Fable 5.1 to an older model and starts behaving differently even with an identical prompt; (3) an SRE reports it is “fine on average” yet users occasionally wait 9 seconds and the monthly bill is a surprise; and (4) a nightly run of 10,000 extraction jobs keeps hitting rate limits.

Expert reasoning trace.

  1. Problem 1 — filling window, keep dialogue. Bulky, stale tool results with the narrative preserved → context editing, not compaction (which would summarise the dialogue) and not a bigger model. Additionally, delegate the search-heavy work to subagents so raw results never enter the main window.

  2. Problem 2 — fallback behaviour change. Fable 5.1 thinking blocks are readable only by that model or newer, so the older fallback silently drops them. The design must test the degraded path and not assume parity; window size and key expiry are red herrings.

  3. Problem 3 — hidden tail and cost. Add p95/p99 latency (the 9-second tail averages hide) and cost per task (the surprise bill). Reject “total request count” and “model name” as non-diagnostic.

  4. Problem 4 — rate limits on bulk. 10,000 latency-tolerant jobs belong on the Batch API (50% off, ≤24h), which eases rate-limit pressure. Reject real-time max concurrency (full price, limit risk) and one giant request (will not fit).

Exam-correct decision: context editing + subagent isolation for the window; a tested degraded path for the Fable 5.1 fallback; p95/p99 and cost-per-task monitoring; and the Batch API for bulk. Every rejected option is a named trap (bigger-model, parity-assumption, average-only monitoring, real-time bulk).


Exam traps in this domain

TrapWhy it is wrong
“Use a bigger model” to fix a filling contextThe fix is context editing/compaction/subagents/memory
Claim Haiku 4.5 has a 1M windowHaiku 4.5 is 200k
Keep must-survive facts only in conversationCompaction/editing are lossy; use memory tool or DB
Use compaction to drop bulky tool resultsThat is context editing’s job; compaction summarises the narrative
Edit old turns to tidy a Fable 5.1 transcriptInvalidates later thinking blocks; harness must be append-only
Fall back from Fable 5.1 assuming parityThinking blocks are dropped; behaviour changes
Retry a 400/413 with backoffNon-retryable; fix the request/size
Retry without jitter or ignoring retry-afterCauses thundering-herd; respect the header
Monitor only average latencyp95 catches the tail; averages hide it
Return nothing on a partial failureReturn partial results with provenance of the gap
“Bigger model” for a cost-rising-every-turn sessionContext editing removes the biggest term; caching cheapens the stable term
Ignore ITPM/RPM when sizing real-time throughputThe token limit often binds first; compute the effective ceiling
Use real-time high concurrency for 10k overnight jobsThe Batch API (50% off, ≤24h) fits and eases rate-limit pressure
Skip timeouts on tool/subagent boundariesA hung call stalls the whole agent; set boundary timeouts
Use Fable 5.1 for a ZDR-required workloadFable 5.1 has 30-day retention and no ZDR; exclude it

Practice questions

Q1 · A long-running agent's answers degrade after many turns as the context fills with large tool outputs the agent no longer needs. What is the BEST fix? (Select one)

A. Switch to a model with a bigger context window. B. Use context editing to clear stale tool results while preserving the conversation narrative. C. Restart the conversation from scratch each time. D. Increase max_tokens.

Answer: B. Clearing bulky, no-longer-needed tool results is exactly context editing. A bigger window (A) delays the problem, restarting (C) loses state, and max_tokens (D) is unrelated.

Q2 · A conversation itself has grown very long but the thread must be preserved for the agent to stay coherent. Which mechanism fits? (Select one)

A. Context editing (clear tool results). B. Compaction: summarise the conversation server-side, preserving the narrative in condensed form. C. Delete the oldest turns and hope. D. Move to Haiku 4.5 for its larger window.

Answer: B. Compaction condenses the narrative when the conversation is the thing that is too long. Context editing (A) targets tool results, deleting turns (C) loses the thread (and breaks Fable 5.1 append-only), and Haiku 4.5 (D) has a smaller window.

Q3 · A customer preference must be available in a session next week. Where should it live? (Select one)

A. Conversation history. B. The memory tool or an external database — durable state that survives compaction and new sessions. C. A thinking block. D. The system prompt of the current session.

Answer: B. Durable, cross-session facts need the memory tool or external state. Conversation history (A) and thinking blocks (C) are lossy/ephemeral, and a single session’s system prompt (D) does not persist.

Q4 · On Claude Fable 5.1, a harness edits earlier turns to remove noise and later responses become inconsistent. What is the cause and fix? (Select one)

A. The model is faulty; open a ticket. B. Editing earlier turns invalidates later thinking blocks; make the harness append-only and reclaim space with context editing/compaction instead. C. Increase the context window. D. Turn thinking off.

Answer: B. Fable 5.1 is append-only; editing turns breaks thinking-block binding. The fix is an append-only harness with server-side trimming. It is not a bug (A), window size (C) is unrelated, and thinking cannot be turned off on Fable 5.1 (D).

Q5 · A system falls back from Fable 5.1 to an older model during an outage and behaviour changes unexpectedly. What is the MOST likely reason? (Select one)

A. The older model has a bigger window. B. Thinking blocks produced by Fable 5.1 are readable only by that model or newer, so the older fallback silently drops them, changing behaviour. C. The API key expired. D. The prompt cache was cold.

Answer: B. Thinking-block binding means older fallbacks drop the thinking, altering behaviour. Window size (A) does not cause this, key expiry (C) would error not change behaviour, and a cold cache (D) affects cost/latency not correctness.

Q6 · Which TWO errors should be retried with exponential backoff and jitter? (Select two)

A. 429 rate_limit. B. 400 invalid_request. C. 529 overloaded. D. 401 authentication. E. 413 request_too_large.

Answer: A and C. 429 and 529 are transient and retryable with backoff (respect retry-after). 400 (B), 401 (D) and 413 (E) are client-side and must be fixed, not retried.

Q7 · A stakeholder says 'just use Haiku 4.5 as the fallback; it has the same 1M window'. What is the correct correction? (Select one)

A. They are right. B. Haiku 4.5’s context window is 200k, not 1M, so large inputs may not fit; the fallback path must be validated. C. Haiku 4.5 has a 2M window. D. Haiku 4.5 cannot be used as a fallback at all.

Answer: B. Haiku 4.5 is 200k. Large-context requests that fit Fable/Opus/Sonnet’s 1M may overflow Haiku. It is not 1M (A), not 2M (C), and can be a fallback if inputs fit (D is too absolute).

Q8 · To improve reliability and cost under load, an architect wants a high cache-hit rate. Which arrangement achieves this? (Select one)

A. Put the variable user input first and the stable system prompt last. B. Put the stable system prompt, tools and reference docs first with cache_control on the last stable block; keep the variable task after. C. Disable caching to avoid stale reads. D. Randomise the prompt order each call.

Answer: B. Caching needs stable content first with a cache boundary; the variable task follows. Reversed order (A) prevents hits, disabling caching (C) raises cost, and randomising (D) destroys the cacheable prefix.

Q9 · A latency-tolerant batch of 10,000 extraction jobs must be processed cheaply. What is the BEST choice? (Select one)

A. Real-time Messages API calls with high concurrency. B. The Message Batches API (50% discount, results within 24h) for latency-tolerant workloads. C. One giant single request. D. Haiku 4.5 in a tight synchronous loop.

Answer: B. The Batch API is designed for latency-tolerant bulk work at 50% off. Real-time high concurrency (A) risks rate limits and costs more, one request (C) will not fit, and a synchronous loop (D) is slow and rate-limited.

Q10 · A subagent must read 50 large files to answer one sub-question. How does delegating this protect reliability? (Select one)

A. It does not; put all 50 files in the coordinator’s context. B. The subagent’s isolated context holds the raw files; only the distilled result returns to the coordinator, keeping the main window small. C. It doubles the cost with no benefit. D. It removes the need for error handling.

Answer: B. Subagent isolation keeps bulky raw material out of the main window and returns only the result. Loading all files into the coordinator (A) bloats context, and B is a real benefit not pure cost (C); error handling still matters (D).

Q11 · An SRE monitors only average latency and misses that some users see 8-second responses. What metric should be added? (Select one)

A. Total request count. B. p95 (and p99) latency, which captures the tail that averages hide. C. The number of tools. D. Model name.

Answer: B. p95/p99 expose the tail experience averages mask. Request count (A), tool count (C) and model name (D) do not reveal tail latency.

Q13 · At turn 30 an agent carries 20k stable tokens, 60k history, 400k accumulated tool results and 10k thinking, and per-turn input cost is rising. Which change reduces cost MOST directly while keeping the dialogue? (Select one)

A. Switch to Haiku 4.5 for a bigger window. B. Use context editing to clear the 400k of stale tool results (the dominant term), dropping this-turn input from ~490k to ~90k, while preserving the narrative. C. Increase max_tokens. D. Disable prompt caching.

Answer: B. The tool-result term dominates; clearing it via context editing removes most of the cost while keeping the dialogue. Haiku 4.5 (A) has a smaller 200k window, max_tokens (C) is output not input, and disabling caching (D) raises cost.

Q14 · A ZDR (zero-data-retention) requirement applies to a workload, and an architect proposes Fable 5.1 for its reasoning. What is the correct critique? (Select one)

A. Fable 5.1 is fine; all models support ZDR. B. Fable 5.1 requires 30-day retention and is not ZDR, so it must be excluded for a ZDR-required workload; choose a compliant model/configuration. C. Enable budget_tokens to turn on ZDR. D. ZDR is only about the context window size.

Answer: B. Fable 5.1 has 30-day retention and no ZDR, disqualifying it where ZDR is required. All models do not support ZDR (A), budget_tokens is unrelated (C), and ZDR is a data-retention property, not window size (D).

Q15 · A tier allows 4,000 RPM and 400,000 input tokens/min; each request uses ~10,000 input tokens. What is the effective real-time throughput ceiling and the implication for 10,000 latency-tolerant jobs? (Select one)

A. 4,000 rpm; run them all in real time. B. ~40 rpm (the token limit binds first: 400,000 / 10,000), so real-time would be slow and limit-prone; use the Batch API (50% off, ≤24h) for the bulk run. C. Unlimited, because tokens do not count. D. 400,000 rpm.

Answer: B. ITPM binds first at ~40 rpm, making real-time bulk slow and risky; the Batch API fits. RPM alone (A) ignores the token limit, tokens always count (C), and D confuses tokens with requests.

Q16 · On Fable 5.1, an architect wants to reclaim window space by rewriting several early turns into a shorter summary in place. Why is this wrong and what is correct? (Select one)

A. It is correct and the cheapest option. B. Rewriting turns breaks Fable 5.1’s append-only history and invalidates later thinking blocks; instead reclaim space with server-side context editing/compaction and keep the harness append-only. C. It is fine if you also lower max_tokens. D. Switch off thinking to allow the rewrite.

Answer: B. Fable 5.1 is append-only; editing turns invalidates thinking blocks. Server-side editing/compaction reclaim space without breaking the chain. It is not correct (A, C), and thinking cannot be disabled on Fable 5.1 (D).

Q17 · A subagent gathering data times out midway. What should the system return to the coordinator? (Select one)

A. Nothing, to be safe. B. A structured partial result with provenance of what is missing (and a retryable timeout error), so the coordinator can retry, proceed with a quorum noting the gap, or escalate. C. A false ‘complete’ answer using only what arrived. D. A bare ‘error’ string.

Answer: B. Partial results plus gap provenance enable an explicit decision. Returning nothing (A) wastes work, a silent false-complete (C) is #7, and a bare error string (D) is #6.

Q18 · An SRE dashboard shows only average latency and total request count, and the team is surprised by both slow tail responses and cost. Which TWO metrics should be added FIRST? (Select two)

A. p95/p99 latency to capture the tail averages hide. B. Cost per task, the unit economics that decide viability. C. The number of tools per agent. D. The model’s name. E. Total prompt character count.

Answer: A and B. p95/p99 exposes the tail and cost-per-task surfaces the economics the team is missing. Tool count (C), model name (D) and character totals (E) do not reveal tail latency or cost.

Key takeaways

  • The context window is a shared budget consumed by system, tools, history, tool results and thinking; keep the working set small.
  • Context editing clears stale tool results; compaction summarises the narrative — neither preserves everything, so durable facts go to the memory tool or a database.
  • Subagent isolation keeps bulky raw material out of the main window; put stable content first and cache it for ~10× cheaper reads and lower latency.
  • Do the arithmetic: the accumulated-tool-results term usually dominates a filling window, so context editing (or subagent isolation) — not a bigger model — is the cost-effective fix.
  • Fable 5.1 is append-only (never edit old turns), thinking is always on, forced tool_choice is a 400, it has 30-day retention/no ZDR, and it is not Priority Tier — design around all of these.
  • Retry only transient errors (429/5xx/529) with backoff, jitter and retry-after; only retry idempotent writes.
  • Fallback is not free — from Fable 5.1 thinking blocks are dropped; Haiku 4.5 is 200k, not 1M; validate the degraded path.
  • Size real-time throughput against ITPM/RPM (the token limit often binds first); use the Batch API for latency-tolerant bulk; set boundary timeouts and return partial results with provenance.
  • Monitor p95/p99 latency, error rate by category, cache-hit rate and cost per task; gate irreversible actions on humans.

Last updated Sep 18, 2026