AI Cert Prep
Type to search documentation.

Domains

D3 · Integration (incl. RAG)

RAG pipeline design end to end, chunking and embeddings, retrieval and reranking, grounding and citations, retrieval evaluation and debugging, tool least-privilege and authz, MCP vs API selection, observability, and enterprise integration patterns.

This is the largest domain on the exam – 19%, roughly 12 of 63 items. It is dominated by RAG architecture but also covers tool/agent capability governance, identity and authorization, integration-mechanism selection, observability at scale, and enterprise integration patterns. The recurring judgement being tested: when an answer is confidently wrong, suspect retrieval and indexing before the model; and when an agent has many tools, apply least privilege by removing capabilities, not by logging or confirming them.

Learning objectives

By the end of this page you should be able to:

  1. Design a RAG pipeline stage by stage: ingestion → chunking → embedding → indexing → retrieval → rerank → grounding.
  2. Match a chunking strategy to the data’s shape.
  3. Choose dense vs sparse vs hybrid retrieval and an index/vector store with metadata filters.
  4. Improve retrieval with reranking, query rewriting, HyDE, multi-query, MMR.
  5. Evaluate retrieval (recall@k, MRR, faithfulness, answer relevance) and debug confident-but-wrong answers.
  6. Decide between RAG, long context (1M) and fine-tuning; design agentic RAG.
  7. Run capability-bloat / least-privilege analysis and close authn/authz gaps.
  8. Choose integration mechanisms (MCP vs API/CLI vs agent-to-agent), design observability at scale, and apply enterprise integration patterns (queues, webhooks, idempotency, batch, freshness).

3.1 The RAG pipeline, stage by stage

Retrieval-Augmented Generation grounds the model in your data by retrieving relevant passages and placing them in context at query time. Learn the pipeline as a fixed skeleton; every design decision slots into one stage.

text
INGESTION INDEX-BUILD (offline) QUERY-TIME (online)
┌──────────┐ ┌───────────────────────────────┐ ┌───────────────────────────────────┐
│ sources │ │ clean → chunk → embed → │ │ query → (rewrite/HyDE/multi-query) │
│ (docs, │──▶│ write vectors + metadata to │ │ → retrieve top-k (dense+sparse) │
│ DB, web)│ │ vector store / search index │ │ → rerank → assemble context │
└──────────┘ └───────────────────────────────┘ │ → Claude generates grounded answer │
│ freshness / refresh pipeline ▲ │ → cite sources → validate │
└──────────────────────────────┘ └───────────────────────────────────┘
StageResponsibilityPrimary failure mode
IngestionPull, clean, normalise, de-duplicate source dataStale/duplicate content; lost structure
ChunkingSplit into retrievable unitsChunks too large (dilute) or too small (fragment)
EmbeddingTurn chunks into vectorsWrong/mismatched embedding model
IndexingStore vectors + metadata for searchNot re-indexed after refresh → stale hits
RetrievalFetch candidate chunks for a queryLow recall; wrong ranking
RerankReorder candidates by true relevanceSkipped → best passage buried below top-k
GroundingConstrain the answer to retrieved text + citeModel answers from parametric memory, not context

3.2 Chunking strategies matched to data shape

Chunking is the highest-leverage RAG decision. The right strategy depends on the structure of the source.

StrategyHow it splitsBest forWeakness
Fixed-sizeN tokens with overlapUniform prose, quick baselineCuts mid-sentence/mid-idea
RecursiveSplit on paragraph → sentence → token boundariesGeneral text with some structureStill boundary-blind to meaning
SemanticSplit where embedding similarity dropsTopically dense text where ideas shiftMore compute to build
Structural / document-awareSplit on headings, sections, table rows, code blocksManuals, contracts, Markdown, HTML, codeRequires a parser per format
Parent–childRetrieve small child chunks, return larger parent for contextPrecise matching + rich context to the modelMore storage/bookkeeping
Late chunkingEmbed the whole doc, then pool per-chunk from token embeddingsLong docs where cross-chunk context mattersNeeds long-context embedding support
text
Document shape?
├─ Structured (headings/tables/code) → structural / document-aware chunking
├─ Long, cross-referential prose → parent–child or late chunking
├─ Topically shifting prose → semantic chunking
└─ Uniform prose, need a baseline → recursive (fall back to fixed)

Exam signal

“Answers cite the wrong section / lose the surrounding context” → the chunks are too small or boundary-blind; move to parent–child or structural chunking. “Contracts / manuals / code” in the stem → document-aware chunking, not fixed-size.


3.3 Embeddings and indexing

Dense vs sparse vs hybrid

Retrieval typeSignalStrong atWeak at
Dense (embeddings)semantic similarityparaphrase, synonyms, conceptsexact IDs, rare tokens, codes
Sparse (BM25 / keyword)lexical overlapexact terms, part numbers, namessynonyms, paraphrase
Hybrid (dense + sparse, fused)both, score-fusedmost enterprise corporaslightly more infra

Most production RAG uses hybrid retrieval (e.g. reciprocal-rank fusion of BM25 and dense) because real queries mix concepts and exact identifiers.

Vector store and metadata filters

Choice driverGuidance
Scale (billions of vectors)Purpose-built vector DB or managed service
Already on a cloudUse its managed vector/search offering (IAM, residency inherited)
Need lexical + vector in oneA search engine with hybrid support
FilteringStore metadata (tenant, doc type, ACL, date) and filter before/along vector search

Metadata filters are also a security control: filtering by tenant_id and per-user ACL at retrieval time is how you prevent cross-tenant/permission leakage (see 3.9).


3.4 Retrieval and re-ranking techniques

TechniqueWhat it doesWhen to add it
top-kfetch the k nearest candidatesalways; tune k for recall vs noise
MMR (max marginal relevance)diversify results, reduce redundancynear-duplicate chunks crowd top-k
Rerankinga cross-encoder re-scores candidates for true relevanceprecision matters; retrieve wide, rerank to few
Query rewritingrephrase/expand the queryconversational or underspecified queries
HyDEgenerate a hypothetical answer, embed that to retrievesparse queries where the answer’s vocabulary differs from the question’s
Multi-queryissue several query variants, union resultsrecall-critical retrieval

A common high-precision recipe: retrieve wide (k=50, hybrid) → rerank → keep top 5–8 → ground. Reranking is often the single biggest precision win.

text
query ─▶ [rewrite / multi-query] ─▶ hybrid retrieve (k=50)
│
▼
reranker ─▶ top 6 ─▶ context ─▶ Claude

3.5 Grounding and citations

Grounding means the answer is constrained to the retrieved passages, and every claim is traceable.

  • Instruct the model to answer only from the provided context and to say when the context is insufficient (never invent).
  • Use the Citations feature / structured references so each claim links to its source chunk and location.
  • Return insufficient context rather than a parametric-memory guess when retrieval fails — a grounded system prefers “I don’t have that” to a confident fabrication.
json
{
"system": "Answer ONLY from <context>. Cite the source id for each claim. If the context does not contain the answer, say so.",
"messages": [
{"role": "user", "content": "<context>{retrieved_chunks_with_ids}</context>\n\nQuestion: What is the termination notice period?"}
]
}

3.6 Evaluating retrieval

You cannot fix what you do not measure. Retrieval and generation are evaluated separately so you know which stage failed.

MetricMeasuresStage
Recall@kDid the relevant chunk make it into the top-k?Retrieval
MRRHow high did the first relevant chunk rank?Retrieval / rerank
Precision@kHow much of top-k is actually relevant?Retrieval / rerank
Faithfulness / groundednessIs the answer supported by retrieved context (no fabrication)?Generation
Answer relevanceDoes the answer address the question?Generation

Exam signal

If recall@k is high but answers are wrong, the failure is in generation/grounding (or reranking), not retrieval. If recall@k is low, fix chunking/embeddings/retrieval first. Measuring only end-to-end accuracy hides which stage broke — an aggregate-metric trap.


3.7 Debugging confident-but-wrong answers after a document refresh

This is a signature exam scenario. A document is updated; the assistant keeps giving the old answer, confidently. The instinct to blame the prompt or the model is wrong.

  1. Suspect the index first. Was the refreshed document re-chunked, re-embedded and re-indexed? A refresh that updates the source store but not the vector index leaves stale vectors that retrieval faithfully returns.

  2. Inspect what was retrieved. Log the retrieved chunk IDs and text for the failing query. If they are the old content, it’s an indexing/freshness bug, not a model bug.

  3. Check chunk/version metadata. Stale updated_at or a missing re-index job confirms it.

  4. Only then examine grounding/prompt. If retrieval returned the new content but the answer used the old, the grounding instruction or reranking is at fault.

text
Confident-but-wrong after a refresh?
1. What did retrieval return? ──stale── ▶ re-chunk/re-embed/re-index (fix freshness pipeline)
2. Returned fresh content? ──yes──── ▶ check grounding instruction / reranking
3. Neither? ────────── ▶ check embedding-model mismatch / query rewrite

The wrong instinct

Distractors will suggest “rewrite the system prompt”, “raise effort”, or “switch to a bigger model”. None fix stale retrieval. The correct first move is to inspect and repair the retrieval/indexing layer.


3.8 RAG vs long context vs fine-tuning

ApproachBest whenCost/ops profileFails when
RAGCorpus is large, changes often, needs citations/freshnessRetrieval infra + refresh pipelineRetrieval quality is poor
Long context (1M)Small/bounded corpus fits the window; simplicity valuedHigh per-call token cost; no infraCorpus too big/costly; needle degradation
Fine-tuningStable domain style/format/behaviour to bake inTraining + retraining cadenceFacts change often (retrain infeasible)
text
Does the knowledge change frequently or need citations? ── yes ─▶ RAG
Is the corpus small, stable, and fits comfortably in context? ── yes ─▶ long context
Is it about behaviour/format/style rather than changing facts? ── yes ─▶ fine-tuning

These are not mutually exclusive: fine-tune for style, RAG for facts, long context for a bounded session.


3.9 Agentic RAG, tool capability-bloat and least privilege

Agentic RAG lets the model decide when and what to retrieve, issue follow-up queries, and combine sources — more powerful than one-shot retrieval, but with more cost and failure surface. Use it when queries are multi-hop or exploratory; use plain RAG when a single retrieval suffices.

Capability-bloat and least privilege

An agent with too many tools (18 vs a recommended 4–5) is slower, more error-prone, and more dangerous. The correct remediation for an unneeded destructive tool is to remove it, not to log its use or add a confirmation prompt.

SymptomWrong fixRight fix
Agent has refund, delete_account, issue_credit it never legitimately needsLog the calls; add “are you sure?”Remove the tools from the agent’s allowlist (least privilege)
18 tools, model picks wrong onesLonger prompt describing eachCut to 4–5; use tool search + defer_loading for large catalogues
Occasional dangerous actionPrompt “never do X”Programmatic permission hook denies X

Least privilege is removal, not observation

Logging a dangerous capability or confirming it still leaves the capability present and reachable (excessive agency). The exam’s correct answer removes unneeded tools so the agent cannot invoke them at all.

Authn / authz gap analysis

Tools act on real systems, so identity and permissions must propagate to the tool call — the agent must act as the user, with the user’s permissions, not as an omnipotent service account.

GapRiskControl
Agent uses one service account for all usersOne user reaches another’s dataPropagate end-user identity; per-user scoping
Tool has no per-user permission checkPrivilege escalationEnforce ACLs inside the tool, not just in the prompt
Remote MCP server unauthenticatedAnyone can call powerful toolsOAuth 2.1 on the remote MCP server
Retrieval ignores ACLsCross-tenant leakageMetadata/ACL filter at retrieval time (3.3)

3.10 Choosing the integration mechanism

MechanismUse whenNotes
MCPReusable tools/resources shared across many agents/clients; standard protocolJSON-RPC 2.0; stdio local / Streamable HTTP remote; OAuth 2.1 for remote; MCP connector lets Messages API call remote servers
Direct API / CLI toolA one-off or app-specific capability; tightest control/latencyDefine as a tool in the request; no protocol overhead
Agent-to-agentDecompose across specialised agents with isolated contextCoordinator/subagent; higher cost/latency
text
Will many clients/agents reuse this capability? ── yes ─▶ MCP server (OAuth if remote)
One app, tight control/latency? ── yes ─▶ direct API/CLI tool
Need a specialised agent with its own context? ── yes ─▶ agent-to-agent (subagent)

Progressive discovery vs monolithic context

Do not load every tool and document into context up front. Use progressive discovery: tool search + defer_loading: true for large tool catalogues, and Skills loaded on demand. A monolithic context is expensive, cache-hostile, and degrades tool-selection accuracy.


3.11 Observability at scale

At production scale you cannot debug what you cannot trace. Instrument every request end to end.

SignalWhat to capture
TracesFull span tree: retrieval, rerank, model call, tool calls, validation
Correlation IDsOne ID threaded through app → model → tools → downstream systems
Token/cost telemetryInput/output/thinking tokens and cost per request, per segment, per model
Retrieval telemetryQuery, retrieved chunk IDs, scores, rerank order
Errors & stop reasonsstop_reason, error codes, retries, fallbacks, degraded-mode flags
text
[req id: 9f3a] ── app ──▶ retrieve(k=50) ──▶ rerank(6) ──▶ opus-5 ──▶ tool:lookup ──▶ validate ──▶ resp
correlation id 9f3a threads through every span; token+cost tagged per span

Per-segment cost/latency telemetry is what lets you route, cache and optimise (D4) with evidence rather than guesses.


3.12 Enterprise integration patterns and data freshness

PatternPurposeClaude-specific note
QueueDecouple bursty producers from rate-limited model callsSmooths load against RPM/ITPM tiers
WebhookReact to external events (ticket created, doc updated)Trigger ingestion/refresh and agent runs
IdempotencySafe retries without duplicate side effectsIdempotency keys on tool actions; critical with backoff retries
BatchLatency-tolerant bulk workMessage Batches API: 50% discount, results within 24 h
Refresh / freshnessKeep the index currentEvent- or schedule-driven re-chunk/re-embed/re-index (ties to 3.7)

Exam signal

“Retried request charged/acted twice” → idempotency keys. “Thousands of documents to classify overnight, cost matters” → Batch API. “Document changed but the answer didn’t” → freshness/refresh pipeline and re-indexing.


3.13 A concrete RAG pipeline configuration

Item writers reward candidates who can read a config and predict its behaviour. A defensible enterprise baseline:

yaml
ingestion:
sources: [confluence, s3_pdfs, postgres_kb]
dedup: content_hash
pii_scrub: true
chunking:
strategy: structural # split on headings/sections
fallback: recursive
target_tokens: 400
overlap_tokens: 60
parent_child: true # match child (~400), return parent (~1500)
embedding:
model: text-embedding-3-large # keep query + index model identical
dimensions: 1024
index:
store: managed_vector_db
metadata: [tenant_id, doc_type, acl, updated_at, source_id]
filters_before_search: [tenant_id, acl] # security + precision
retrieval:
mode: hybrid # BM25 + dense, RRF fusion
k: 50
rerank:
model: cross_encoder_reranker
keep_top: 6
query_transform: [rewrite, multi_query] # for conversational/sparse queries
generation:
model: claude-sonnet-5
grounding: answer_only_from_context
citations: true
on_insufficient_context: "say so; do not use parametric memory"
refresh:
trigger: [webhook_on_update, nightly_schedule]
action: re_chunk_re_embed_re_index
KnobEffect if too lowEffect if too high
target_tokensfragments ideas; loses contextdilutes relevance; buries the answer
overlap_tokensboundary facts loststorage/cost bloat, duplicate hits
k (pre-rerank)low recallnoise; slower rerank
keep_topbest passage may be droppedcontext bloat, higher cost, needle dilution

Exam signal

“Retrieve wide (k≈50 hybrid) → rerank → keep 6 → ground with citations” is the canonical high-precision recipe. Distractors that skip reranking, use dense-only, or raise keep_top to 50 are the wrong answers.


3.14 Retrieval evaluation with worked numbers

You must be able to compute the metrics, not just name them.

Setup. 5 queries; for each we know the single relevant chunk and where it ranked in the retrieved list:

QueryRank of relevant chunkIn top-3?Reciprocal rank
Q11yes1/1 = 1.00
Q24no1/4 = 0.25
Q32yes1/2 = 0.50
Q4(not retrieved)no0
Q51yes1/1 = 1.00
text
recall@3 = (#queries whose relevant chunk is in top-3) / total = 3/5 = 0.60
MRR = mean reciprocal rank = (1.00 + 0.25 + 0.50 + 0 + 1.00) / 5 = 0.55

Now suppose adding a reranker moves Q2’s relevant chunk to rank 2 and retrieves Q4’s at rank 3:

text
recall@3 → 5/5 = 1.00 (both now in top-3)
MRR → (1.00 + 0.50 + 0.50 + 0.333 + 1.00)/5 = 0.667

The reranker lifted recall@3 from 0.60 → 1.00 and MRR from 0.55 → 0.67 — the single biggest precision/ranking win, without touching chunking or embeddings.

MetricFormulaReads as
recall@krelevant-in-top-k ÷ total queries“did we retrieve the answer at all?”
MRRmean(1 ÷ rank of first relevant)“how high did it rank?”
precision@krelevant-in-top-k ÷ k“how clean is the top-k?”
faithfulnesssupported claims ÷ total claims“did the answer stay grounded?”

Interpreting the numbers

High recall but low faithfulness ⇒ retrieval is fine, generation/grounding is broken. Low recall ⇒ fix chunking/embeddings/retrieval first. Reporting only end-to-end accuracy hides which of these is true — the aggregate-metric trap.


3.15 Multi-tenant isolation and ACL-aware retrieval

In a shared index, retrieval is a security boundary. A query must never surface another tenant’s or an unauthorised user’s chunk.

text
user (tenant_42, role: support) asks a question
│ attach identity: {tenant_id: 42, acl_groups: [support]}
▼
filter BEFORE/ALONGSIDE vector search:
WHERE tenant_id = 42 AND acl IN ('public','support')
▼
hybrid retrieve → rerank → ground (only authorised chunks ever enter context)
GapLeakControl
No tenant_id filterCross-tenant data exposureMandatory pre-filter on tenant_id
ACL applied only in the promptModel can be talked past itEnforce ACL at the retrieval layer, not in prose
Shared service account for retrievalUser sees data they can’t accessPropagate end-user identity into the query filter
Rerank/cache ignores tenantCached cross-tenant hitKey caches by tenant + ACL

Retrieval-time filtering is the control

Filtering after generation (“the model won’t mention it”) is not a control — the unauthorised chunk already entered context and can leak. The exam’s correct answer filters by tenant_id/ACL before the vector search.


3.16 Scenario walkthrough: a leaking, stale, over-privileged support RAG agent

Scenario. A B2B SaaS support agent serves 300 tenants from one shared vector index. Three complaints arrive: (1) a tenant occasionally sees another tenant’s runbook in an answer; (2) after customers update their docs, the agent still quotes yesterday’s version, confidently; (3) the agent has 22 tools including delete_ticket and refund it should never call. Retrieval is dense-only, k=8, no reranker.

Expert reasoning trace.

  1. Isolation first (highest severity). Cross-tenant exposure is a security incident. Add a mandatory tenant_id (and ACL) pre-filter on every query, key caches by tenant, and propagate end-user identity. Do not rely on prompt wording to keep tenants apart.

  2. Freshness second. Confident-wrong-after-update is stale retrieval: the refresh updated the source store but not the vector index. Inspect retrieved chunk IDs (they’ll be old), then fix the re-chunk/re-embed/re-index pipeline with webhook + nightly triggers.

  3. Least privilege third. Remove delete_ticket, refund, and other unneeded tools from the allowlist — do not merely log or add “are you sure?”. Cut to ~4–5 tools; use tool search + defer_loading if the legitimate catalogue is large.

  4. Then quality. Dense-only + k=8 + no rerank underperforms on exact identifiers and precision. Move to hybrid retrieval, k=50, rerank, keep top 6, and measure recall@k and faithfulness per tenant segment.

  5. Close the loop. Instrument correlation IDs and per-segment retrieval telemetry so the next incident is reconstructable.

Why the tempting alternatives are wrong: “add a system-prompt rule not to reveal other tenants” is prompt-as-enforcement and leaves the leak; “switch to a bigger model” fixes neither isolation nor freshness; “log the dangerous tools” leaves excessive agency; “raise k to 500” adds noise instead of adding a reranker.


3.17 Common misconceptions

MisconceptionRealityWhy it matters on the exam
“Confident-but-wrong means the model is bad.”After a data change it usually means stale retrieval/indexing.Inspect retrieval first; prompt/model fixes are distractors.
“Dense embeddings retrieve everything.”Dense is weak on exact IDs/codes; hybrid adds sparse.Part-number/SKU stems require hybrid.
“A bigger context window replaces RAG.”It costs more per call, can’t cite, and degrades on the needle.Long-context-stuffing is the wrong answer for large/changing corpora.
“Fine-tuning is how you add knowledge.”Fine-tuning bakes behaviour/style; changing facts need RAG.Weekly-changing-facts stems reject fine-tuning.
“Logging a dangerous tool makes it safe.”The capability is still reachable — excessive agency.Least privilege means removal, not observation.
“Prompt rules keep tenants isolated.”Isolation must be enforced at the retrieval filter.Retrieval-time ACL filtering is the correct control.
“MCP servers are authenticated by default.”Remote MCP needs OAuth 2.1 + per-user checks.Unauthenticated remote MCP is a security trap.
“Retrying is always safe.”Non-idempotent actions double up; add idempotency keys.Double-charge stems test idempotency.

Exam traps in this domain

TrapWhy it is wrong
Blaming the prompt/model for confident-wrong answers after a refreshThe cause is usually stale retrieval/indexing; inspect what was retrieved first
Fixing an unneeded dangerous tool by logging or confirming itLeaves excessive agency; least privilege means removing the tool
Giving an agent 18 tools “for flexibility”Slower, error-prone; cut to 4–5, use tool search + defer_loading
Using dense-only retrieval for part numbers / exact IDsDense misses exact tokens; use sparse/hybrid
Measuring only end-to-end accuracyHides whether retrieval or generation failed; measure recall@k and faithfulness separately
Long-context stuffing a large, changing corpusCostly, no citations, needle degradation; use RAG
Fine-tuning to inject frequently-changing factsRetraining cadence infeasible; use RAG
Skipping rerankingBest passage buried below top-k; precision suffers
One service account for all users’ tool callsCross-user data exposure; propagate identity + per-user ACLs
Unauthenticated remote MCP serverAnyone can invoke powerful tools; require OAuth 2.1
Retrying non-idempotent tool actions without keysDuplicate side effects (double refund/charge)
Loading all tools/docs into context up frontExpensive, cache-hostile, worse tool selection; use progressive discovery
Enforcing multi-tenant isolation with a system-prompt ruleIsolation must be a retrieval-time tenant_id/ACL filter, not prose
Filtering unauthorised chunks after generationThe chunk already entered context; filter before the vector search
Raising k to hundreds instead of adding a rerankerAdds noise; reranking is the precision/ranking win
Reporting only end-to-end accuracy for a RAG systemHides whether retrieval or grounding failed; measure recall@k and faithfulness
Using the same embedding model for query and index inconsistentlyQuery/index model mismatch wrecks similarity; keep them identical
Caching retrieval results without keying by tenant/ACLRisks serving a cross-tenant cached hit

Practice questions

Q1 · A policy assistant kept answering with the OLD figure after a policy document was updated last night, and it sounds completely confident. What should the architect investigate FIRST? (Select one)

A. Rewrite the system prompt to be more accurate. B. Inspect what retrieval actually returned; the updated document was likely not re-chunked/re-embedded/re-indexed, so retrieval is serving stale vectors. C. Switch to Opus 5 with xhigh effort. D. Add more few-shot examples.

Answer: B. Confident-but-wrong immediately after a refresh points to stale retrieval/indexing, not the model. The first move is to log the retrieved chunks and confirm whether the refresh updated the vector index. Prompt rewrites (A, D) and a bigger model (C) cannot fix stale retrieval.

Q2 · A support agent has 18 tools, including `delete_account` and `issue_refund`, which its role should never use. What is the correct remediation? (Select one)

A. Keep the tools but log every call for audit. B. Remove the unneeded tools from the agent’s allowlist (least privilege) so it cannot invoke them at all. C. Add a confirmation prompt before those tools run. D. Add a system-prompt sentence forbidding their use.

Answer: B. Least privilege means the capability should not be present. Removing the tools eliminates the excessive agency. Logging (A) and confirmation (C) leave the capability reachable; a prompt rule (D) is prompt-as-enforcement and can be bypassed.

Q3 · Users search a parts catalogue by exact part numbers AND by descriptions. Dense-only retrieval misses many exact-number queries. What is the BEST fix? (Select one)

A. Increase k to 500. B. Use hybrid retrieval (BM25 + dense with rank fusion) so exact identifiers and semantic matches both rank well. C. Switch to a bigger generation model. D. Remove metadata filters.

Answer: B. Dense embeddings are weak on exact tokens/codes; sparse (BM25) handles them, and hybrid fusion covers both query types. A huge k (A) adds noise without fixing lexical matching; a bigger model (C) doesn’t change what is retrieved; removing filters (D) hurts precision and security.

Q4 · Retrieval eval shows recall@10 = 0.95 but faithfulness is low and answers include facts not in the retrieved chunks. Where is the failure and the fix? (Select one)

A. Retrieval; lower k. B. Generation/grounding; strengthen the instruction to answer only from context, add citations, and consider reranking so the best passage is on top. C. Embeddings; change the model. D. Indexing; re-index everything.

Answer: B. High recall means the right chunks are retrieved, so the fault is in generation/grounding — the model is answering from parametric memory. Grounding instructions, citations and reranking address it. The retrieval-side fixes (A, C, D) target a stage that is already performing well.

Q5 · A 2M-document knowledge base changes weekly and answers must cite the exact source clause. Which approach is BEST? (Select one)

A. Fine-tune a model on the corpus each week. B. RAG with hybrid retrieval, reranking and citations, plus a scheduled refresh pipeline. C. Stuff the whole corpus into the 1M context per query. D. Long context plus fine-tuning combined.

Answer: B. Large, frequently-changing, citation-requiring corpora are the canonical RAG case. Weekly fine-tuning (A) is an infeasible retraining cadence for facts. The corpus exceeds the window and would cost too much and lose citations (C). (D) inherits both problems.

Q6 · A contract-QA system chunks contracts every 500 tokens with fixed size. Answers cite the wrong sub-clause and lose surrounding context. Which TWO changes help most? (Select two)

A. Use structural/document-aware chunking that splits on clauses/sections. B. Use parent–child chunking: match on small child chunks but return the larger parent clause for context. C. Switch to dense-only retrieval. D. Increase temperature. E. Remove citations.

Answer: A and B. Contracts are structured, so clause/section-aware chunking preserves boundaries, and parent–child gives precise matching with enough surrounding context. Dense-only (C), temperature (D) and removing citations (E) do not address the chunking problem — the last two make it worse.

Q7 · A remote MCP server exposes powerful tools over Streamable HTTP with no auth. What must the architect add? (Select one)

A. Nothing; MCP is safe by default. B. OAuth 2.1 authentication on the remote MCP server, plus per-user permission checks inside the tools. C. A longer system prompt. D. A higher rate-limit tier.

Answer: B. Remote MCP servers require OAuth 2.1, and tools must enforce per-user permissions so the agent acts with the caller’s authority. MCP is not authenticated by default (A); prompts (C) and rate limits (D) do not address authorization.

Q8 · A nightly job must classify 200,000 documents; latency is not important but cost is. Which mechanism is BEST? (Select one)

A. Real-time synchronous calls in a tight loop. B. The Message Batches API for a 50% discount with results within 24 hours. C. A multi-agent system. D. Fine-tuning.

Answer: B. Latency-tolerant bulk work is exactly the Batch API’s use case (50% discount, results within 24 h). Synchronous loops (A) hit rate limits and cost more. Multi-agent (C) adds cost/complexity; fine-tuning (D) is unrelated to a classification batch.

Q9 · After enabling backoff retries, some refunds are issued twice. What is the correct fix? (Select one)

A. Disable retries entirely. B. Add idempotency keys to the refund tool so retried calls are de-duplicated and produce no duplicate side effect. C. Lower the model temperature. D. Log the duplicates and reconcile later.

Answer: B. Retries on non-idempotent actions cause duplicate side effects; idempotency keys make retries safe. Disabling retries (A) harms resilience. Temperature (C) is irrelevant. Reconciling after the fact (D) still charged customers twice.

Q10 · A single retrieval sometimes isn't enough — some questions need follow-up lookups combining multiple sources. Which design fits, and what is the trade-off? (Select one)

A. Agentic RAG, where the model decides when/what to retrieve and can issue follow-up queries, at the cost of more latency and tokens. B. Stuff everything into context to avoid retrieval. C. Fine-tune on the multi-hop questions. D. Remove reranking to speed things up.

Answer: A. Multi-hop, exploratory queries justify agentic RAG’s model-driven retrieval loop; the trade-off is higher cost and latency versus one-shot RAG. Context stuffing (B) doesn’t scale; fine-tuning (C) can’t hold changing facts; removing reranking (D) hurts precision.

Q11 · A capability will be reused by many agents and clients across the company and must follow a standard protocol. Which integration mechanism is BEST? (Select one)

A. Hard-code it as a per-app CLI tool in each service. B. Build it as an MCP server (with OAuth 2.1 if remote) so many clients reuse it via a standard protocol. C. Implement it as an agent-to-agent handoff only. D. Paste its logic into every system prompt.

Answer: B. Reuse across many clients under a standard protocol is MCP’s purpose. Per-app CLI tools (A) fragment the implementation; agent-to-agent (C) is for specialised context isolation, not shared capability; prompt-pasting (D) is unmaintainable.

Q12 · An agent loads all 40 tools and 30 reference documents into context on every request; cost is high, cache hit rate is low, and tool selection is error-prone. What is the BEST remedy? (Select two)

A. Use tool search with defer_loading: true so only relevant tools are loaded. B. Package reference material as Skills loaded progressively on demand. C. Increase the context window to 1M and keep loading everything. D. Escalate every request to Opus 5. E. Disable prompt caching.

Answer: A and B. Progressive discovery — tool search with deferred loading and on-demand Skills — keeps context lean, restores a stable cache prefix, and improves tool selection. A bigger window (C) still pays for the bloat; escalating models (D) raises cost; disabling caching (E) is the opposite of the fix.

Q13 · A B2B assistant serves 300 tenants from one shared vector index; a tenant occasionally sees another tenant's document in an answer. What is the correct control? (Select one)

A. Add a system-prompt rule telling the model not to reveal other tenants’ data. B. Apply a mandatory tenant_id (and ACL) filter at retrieval time, before/alongside the vector search, so only authorised chunks ever enter context. C. Filter the answer after generation to remove other tenants’ data. D. Give each tenant a bigger model.

Answer: B. Isolation is a retrieval-time security boundary: filter by tenant_id/ACL before the vector search. A prompt rule (A) is bypassable prompt-as-enforcement; post-generation filtering (C) is too late — the chunk already entered context; a bigger model (D) doesn’t isolate data.

Q14 · Across 5 queries the relevant chunk ranked 1, 4, 2, not-retrieved, 1. What are recall@3 and MRR? (Select one)

A. recall@3 = 1.00; MRR = 1.00. B. recall@3 = 0.60; MRR = 0.55. C. recall@3 = 0.55; MRR = 0.60. D. recall@3 = 0.80; MRR = 0.70.

Answer: B. Three of five relevant chunks are in the top 3 (ranks 1, 2, 1) → recall@3 = 3/5 = 0.60. Reciprocal ranks are 1, 0.25, 0.5, 0, 1 → MRR = 2.75/5 = 0.55. Option A ignores the misses; C swaps the two values; D is arithmetically wrong.

Q15 · recall@10 = 0.62 (low) and answers are frequently missing the needed fact. Where should the architect work FIRST? (Select one)

A. Grounding; tighten the answer-only-from-context instruction. B. Retrieval: fix chunking/embeddings/hybrid and add reranking, because low recall means the relevant chunk often isn’t retrieved at all. C. Add more citations. D. Switch the generation model to Opus 5.

Answer: B. Low recall means the right chunk isn’t reaching the top-k, so the failure is upstream in retrieval. Grounding/citations (A, C) and a bigger generation model (D) can’t help if the answer was never retrieved.

Q16 · A RAG system indexes with `text-embedding-3-large` but a new service queries with a different embedding model. Similarity scores look random. What is the cause and fix? (Select one)

A. The vector store is corrupt; rebuild hardware. B. Query/index embedding-model mismatch; use the identical embedding model for both indexing and querying. C. k is too low; raise it to 1000. D. The generation model is too small.

Answer: B. Embeddings from different models live in different vector spaces, so cross-model similarity is meaningless; query and index must use the same embedding model. It isn’t hardware (A); raising k (C) can’t fix incompatible vectors; the generation model (D) is unrelated to retrieval similarity.

Q17 · An agent retrieves k=8 with dense-only and no reranker; precision is poor and exact part numbers are missed. Which TWO changes give the biggest quality lift? (Select two)

A. Switch to hybrid retrieval (BM25 + dense) so exact identifiers rank. B. Retrieve wide (k≈50) and add a reranker, keeping the top 6. C. Increase temperature. D. Remove citations to speed responses. E. Move to a 1M-token context and stuff everything.

Answer: A and B. Hybrid retrieval fixes exact-identifier misses and retrieve-wide-then-rerank fixes precision — the two canonical levers. Temperature (C) is irrelevant to retrieval; removing citations (D) harms grounding traceability; stuffing context (E) inflates cost without improving ranking.

Q18 · Which TWO signals let you reconstruct a failed multi-step RAG request end to end? (Select two)

A. A correlation ID threaded through app → retrieval → model → tools → downstream. B. Per-span traces capturing retrieved chunk IDs, rerank order, token/cost, and stop_reason. C. Only the final HTTP status code. D. The model’s own assessment that it did fine. E. A daily aggregate request count.

Answer: A and B. A correlation ID plus per-span traces (with retrieval detail and cost) make an incident reconstructable. A status code (C) and a daily count (E) are too coarse; self-assessment (D) is unreliable (self-report anti-pattern).

Key takeaways

  • RAG is a fixed pipeline: ingest → chunk → embed → index → retrieve → rerank → ground → cite; every decision slots into a stage.
  • Match chunking to data shape: document-aware for structured text, parent–child/late for cross-referential prose, semantic for shifting topics.
  • Use hybrid retrieval (dense + sparse) for real corpora; rerank for precision; add rewriting/HyDE/multi-query for recall.
  • Ground answers to retrieved context, cite sources, and prefer “insufficient context” over fabrication.
  • Evaluate retrieval and generation separately (recall@k, MRR vs faithfulness); confident-wrong-after-refresh means inspect retrieval/indexing first.
  • Choose RAG for changing/cited facts, long context for small stable corpora, fine-tuning for behaviour/style.
  • Apply least privilege by removing unneeded (especially destructive) tools; keep agents to ~4–5 tools with tool search for larger catalogues.
  • Propagate user identity to tools, enforce per-user ACLs, and require OAuth 2.1 on remote MCP servers.
  • Instrument traces, correlation IDs and token/cost telemetry; use queues, idempotency keys, webhooks, Batch API and a freshness pipeline for enterprise integration.
  • Compute retrieval metrics: recall@k = relevant-in-top-k ÷ queries; MRR = mean(1÷rank); a reranker typically lifts both the most.
  • In a shared index, retrieval is a security boundary: filter by tenant_id/ACL before the vector search, propagate end-user identity, and key caches by tenant.
  • Keep the query and index embedding models identical; the canonical recipe is retrieve wide (k≈50 hybrid) → rerank → keep ~6 → ground with citations.

Last updated Sep 18, 2026