# D3 · Integration (incl. RAG)

RAG pipeline design end to end, chunking and embeddings, retrieval and reranking, grounding and citations, retrieval evaluation and debugging, tool least-privilege and authz, MCP vs API selection, observability, and enterprise integration patterns.

import { Accordions, AccordionItem, Tabs, TabItem, Steps } from '@prosefly/astro-components';

This is the **largest domain on the exam – 19%, roughly 12 of 63 items**. It is dominated by **RAG architecture** but also covers tool/agent capability governance, identity and authorization, integration-mechanism selection, observability at scale, and enterprise integration patterns. The recurring judgement being tested: when an answer is *confidently wrong*, suspect **retrieval and indexing before the model**; and when an agent has many tools, apply **least privilege** by removing capabilities, not by logging or confirming them.

## Learning objectives

By the end of this page you should be able to:

1. Design a **RAG pipeline** stage by stage: ingestion → chunking → embedding → indexing → retrieval → rerank → grounding.
2. Match a **chunking strategy** to the data's shape.
3. Choose **dense vs sparse vs hybrid** retrieval and an index/vector store with metadata filters.
4. Improve retrieval with **reranking, query rewriting, HyDE, multi-query, MMR**.
5. **Evaluate retrieval** (recall@k, MRR, faithfulness, answer relevance) and debug confident-but-wrong answers.
6. Decide between **RAG, long context (1M) and fine-tuning**; design **agentic RAG**.
7. Run **capability-bloat / least-privilege** analysis and close **authn/authz** gaps.
8. Choose integration mechanisms (**MCP vs API/CLI vs agent-to-agent**), design **observability at scale**, and apply **enterprise integration patterns** (queues, webhooks, idempotency, batch, freshness).

---

## 3.1 The RAG pipeline, stage by stage

Retrieval-Augmented Generation grounds the model in *your* data by retrieving relevant passages and placing them in context at query time. Learn the pipeline as a fixed skeleton; every design decision slots into one stage.

```text
  INGESTION            INDEX-BUILD (offline)                    QUERY-TIME (online)
 ┌──────────┐   ┌───────────────────────────────┐   ┌───────────────────────────────────┐
 │ sources  │   │ clean → chunk → embed →        │   │ query → (rewrite/HyDE/multi-query) │
 │ (docs,   │──▶│ write vectors + metadata to    │   │  → retrieve top-k (dense+sparse)   │
 │  DB, web)│   │ vector store / search index    │   │  → rerank → assemble context       │
 └──────────┘   └───────────────────────────────┘   │  → Claude generates grounded answer │
      │  freshness / refresh pipeline ▲              │  → cite sources → validate          │
      └──────────────────────────────┘              └───────────────────────────────────┘
```

| Stage | Responsibility | Primary failure mode |
| --- | --- | --- |
| Ingestion | Pull, clean, normalise, de-duplicate source data | Stale/duplicate content; lost structure |
| Chunking | Split into retrievable units | Chunks too large (dilute) or too small (fragment) |
| Embedding | Turn chunks into vectors | Wrong/mismatched embedding model |
| Indexing | Store vectors + metadata for search | Not re-indexed after refresh → stale hits |
| Retrieval | Fetch candidate chunks for a query | Low recall; wrong ranking |
| Rerank | Reorder candidates by true relevance | Skipped → best passage buried below top-k |
| Grounding | Constrain the answer to retrieved text + cite | Model answers from parametric memory, not context |

---

## 3.2 Chunking strategies matched to data shape

Chunking is the highest-leverage RAG decision. The right strategy depends on the **structure of the source**.

| Strategy | How it splits | Best for | Weakness |
| --- | --- | --- | --- |
| Fixed-size | N tokens with overlap | Uniform prose, quick baseline | Cuts mid-sentence/mid-idea |
| Recursive | Split on paragraph → sentence → token boundaries | General text with some structure | Still boundary-blind to meaning |
| Semantic | Split where embedding similarity drops | Topically dense text where ideas shift | More compute to build |
| Structural / document-aware | Split on headings, sections, table rows, code blocks | Manuals, contracts, Markdown, HTML, code | Requires a parser per format |
| Parent–child | Retrieve small child chunks, return larger parent for context | Precise matching + rich context to the model | More storage/bookkeeping |
| Late chunking | Embed the whole doc, then pool per-chunk from token embeddings | Long docs where cross-chunk context matters | Needs long-context embedding support |

```text
Document shape?
  ├─ Structured (headings/tables/code) → structural / document-aware chunking
  ├─ Long, cross-referential prose      → parent–child or late chunking
  ├─ Topically shifting prose           → semantic chunking
  └─ Uniform prose, need a baseline     → recursive (fall back to fixed)
```

:::tip[Exam signal]
"Answers cite the wrong section / lose the surrounding context" → the chunks are too small or boundary-blind; move to **parent–child** or **structural** chunking. "Contracts / manuals / code" in the stem → **document-aware** chunking, not fixed-size.
:::

---

## 3.3 Embeddings and indexing

### Dense vs sparse vs hybrid

| Retrieval type | Signal | Strong at | Weak at |
| --- | --- | --- | --- |
| Dense (embeddings) | semantic similarity | paraphrase, synonyms, concepts | exact IDs, rare tokens, codes |
| Sparse (BM25 / keyword) | lexical overlap | exact terms, part numbers, names | synonyms, paraphrase |
| Hybrid (dense + sparse, fused) | both, score-fused | most enterprise corpora | slightly more infra |

Most production RAG uses **hybrid** retrieval (e.g. reciprocal-rank fusion of BM25 and dense) because real queries mix concepts and exact identifiers.

### Vector store and metadata filters

| Choice driver | Guidance |
| --- | --- |
| Scale (billions of vectors) | Purpose-built vector DB or managed service |
| Already on a cloud | Use its managed vector/search offering (IAM, residency inherited) |
| Need lexical + vector in one | A search engine with hybrid support |
| Filtering | Store **metadata** (tenant, doc type, ACL, date) and filter *before/along* vector search |

**Metadata filters are also a security control**: filtering by `tenant_id` and per-user ACL at retrieval time is how you prevent cross-tenant/permission leakage (see 3.9).

---

## 3.4 Retrieval and re-ranking techniques

| Technique | What it does | When to add it |
| --- | --- | --- |
| top-k | fetch the k nearest candidates | always; tune k for recall vs noise |
| MMR (max marginal relevance) | diversify results, reduce redundancy | near-duplicate chunks crowd top-k |
| Reranking | a cross-encoder re-scores candidates for true relevance | precision matters; retrieve wide, rerank to few |
| Query rewriting | rephrase/expand the query | conversational or underspecified queries |
| HyDE | generate a hypothetical answer, embed *that* to retrieve | sparse queries where the answer's vocabulary differs from the question's |
| Multi-query | issue several query variants, union results | recall-critical retrieval |

A common high-precision recipe: **retrieve wide (k=50, hybrid) → rerank → keep top 5–8 → ground**. Reranking is often the single biggest precision win.

```text
query ─▶ [rewrite / multi-query] ─▶ hybrid retrieve (k=50)
                                        │
                                        ▼
                                    reranker  ─▶ top 6 ─▶ context ─▶ Claude
```

---

## 3.5 Grounding and citations

Grounding means the answer is **constrained to the retrieved passages**, and every claim is traceable.

- Instruct the model to answer **only** from the provided context and to say when the context is insufficient (never invent).
- Use the **Citations** feature / structured references so each claim links to its source chunk and location.
- Return `insufficient context` rather than a parametric-memory guess when retrieval fails — a grounded system prefers "I don't have that" to a confident fabrication.

```json
{
  "system": "Answer ONLY from <context>. Cite the source id for each claim. If the context does not contain the answer, say so.",
  "messages": [
    {"role": "user", "content": "<context>{retrieved_chunks_with_ids}</context>\n\nQuestion: What is the termination notice period?"}
  ]
}
```

---

## 3.6 Evaluating retrieval

You cannot fix what you do not measure. Retrieval and generation are evaluated **separately** so you know which stage failed.

| Metric | Measures | Stage |
| --- | --- | --- |
| Recall@k | Did the relevant chunk make it into the top-k? | Retrieval |
| MRR | How high did the first relevant chunk rank? | Retrieval / rerank |
| Precision@k | How much of top-k is actually relevant? | Retrieval / rerank |
| Faithfulness / groundedness | Is the answer supported by retrieved context (no fabrication)? | Generation |
| Answer relevance | Does the answer address the question? | Generation |

:::tip[Exam signal]
If **recall@k is high but answers are wrong**, the failure is in generation/grounding (or reranking), not retrieval. If **recall@k is low**, fix chunking/embeddings/retrieval first. Measuring only end-to-end accuracy hides which stage broke — an aggregate-metric trap.
:::

---

## 3.7 Debugging confident-but-wrong answers after a document refresh

This is a signature exam scenario. A document is updated; the assistant keeps giving the **old** answer, confidently. The instinct to blame the prompt or the model is wrong.

<Steps>

1. **Suspect the index first.** Was the refreshed document re-chunked, re-embedded and re-indexed? A refresh that updates the source store but **not the vector index** leaves stale vectors that retrieval faithfully returns.

2. **Inspect what was retrieved.** Log the retrieved chunk IDs and text for the failing query. If they are the *old* content, it's an indexing/freshness bug, not a model bug.

3. **Check chunk/version metadata.** Stale `updated_at` or a missing re-index job confirms it.

4. **Only then** examine grounding/prompt. If retrieval returned the *new* content but the answer used the old, the grounding instruction or reranking is at fault.

</Steps>

```text
Confident-but-wrong after a refresh?
  1. What did retrieval return?  ──stale── ▶ re-chunk/re-embed/re-index (fix freshness pipeline)
  2. Returned fresh content?     ──yes──── ▶ check grounding instruction / reranking
  3. Neither?                    ────────── ▶ check embedding-model mismatch / query rewrite
```

:::caution[The wrong instinct]
Distractors will suggest "rewrite the system prompt", "raise effort", or "switch to a bigger model". None fix stale retrieval. The correct first move is to **inspect and repair the retrieval/indexing layer**.
:::

---

## 3.8 RAG vs long context vs fine-tuning

| Approach | Best when | Cost/ops profile | Fails when |
| --- | --- | --- | --- |
| RAG | Corpus is large, changes often, needs citations/freshness | Retrieval infra + refresh pipeline | Retrieval quality is poor |
| Long context (1M) | Small/bounded corpus fits the window; simplicity valued | High per-call token cost; no infra | Corpus too big/costly; needle degradation |
| Fine-tuning | Stable domain style/format/behaviour to bake in | Training + retraining cadence | Facts change often (retrain infeasible) |

```text
Does the knowledge change frequently or need citations? ── yes ─▶ RAG
Is the corpus small, stable, and fits comfortably in context? ── yes ─▶ long context
Is it about behaviour/format/style rather than changing facts? ── yes ─▶ fine-tuning
```

These are not mutually exclusive: fine-tune for style, RAG for facts, long context for a bounded session.

---

## 3.9 Agentic RAG, tool capability-bloat and least privilege

**Agentic RAG** lets the model decide *when* and *what* to retrieve, issue follow-up queries, and combine sources — more powerful than one-shot retrieval, but with more cost and failure surface. Use it when queries are multi-hop or exploratory; use plain RAG when a single retrieval suffices.

### Capability-bloat and least privilege

An agent with too many tools (18 vs a recommended 4–5) is slower, more error-prone, and more dangerous. The **correct remediation for an unneeded destructive tool is to remove it**, not to log its use or add a confirmation prompt.

| Symptom | Wrong fix | Right fix |
| --- | --- | --- |
| Agent has `refund`, `delete_account`, `issue_credit` it never legitimately needs | Log the calls; add "are you sure?" | **Remove** the tools from the agent's allowlist (least privilege) |
| 18 tools, model picks wrong ones | Longer prompt describing each | Cut to 4–5; use **tool search + `defer_loading`** for large catalogues |
| Occasional dangerous action | Prompt "never do X" | Programmatic permission hook denies X |

:::caution[Least privilege is removal, not observation]
Logging a dangerous capability or confirming it still leaves the capability present and reachable (excessive agency). The exam's correct answer **removes** unneeded tools so the agent cannot invoke them at all.
:::

### Authn / authz gap analysis

Tools act on real systems, so **identity and permissions must propagate to the tool call** — the agent must act *as the user*, with the user's permissions, not as an omnipotent service account.

| Gap | Risk | Control |
| --- | --- | --- |
| Agent uses one service account for all users | One user reaches another's data | Propagate end-user identity; per-user scoping |
| Tool has no per-user permission check | Privilege escalation | Enforce ACLs inside the tool, not just in the prompt |
| Remote MCP server unauthenticated | Anyone can call powerful tools | **OAuth 2.1** on the remote MCP server |
| Retrieval ignores ACLs | Cross-tenant leakage | Metadata/ACL filter at retrieval time (3.3) |

---

## 3.10 Choosing the integration mechanism

| Mechanism | Use when | Notes |
| --- | --- | --- |
| **MCP** | Reusable tools/resources shared across many agents/clients; standard protocol | JSON-RPC 2.0; `stdio` local / Streamable HTTP remote; OAuth 2.1 for remote; MCP connector lets Messages API call remote servers |
| **Direct API / CLI tool** | A one-off or app-specific capability; tightest control/latency | Define as a `tool` in the request; no protocol overhead |
| **Agent-to-agent** | Decompose across specialised agents with isolated context | Coordinator/subagent; higher cost/latency |

```text
Will many clients/agents reuse this capability?  ── yes ─▶ MCP server (OAuth if remote)
One app, tight control/latency?                  ── yes ─▶ direct API/CLI tool
Need a specialised agent with its own context?   ── yes ─▶ agent-to-agent (subagent)
```

### Progressive discovery vs monolithic context

Do **not** load every tool and document into context up front. Use **progressive discovery**: **tool search + `defer_loading: true`** for large tool catalogues, and **Skills** loaded on demand. A monolithic context is expensive, cache-hostile, and degrades tool-selection accuracy.

---

## 3.11 Observability at scale

At production scale you cannot debug what you cannot trace. Instrument every request end to end.

| Signal | What to capture |
| --- | --- |
| Traces | Full span tree: retrieval, rerank, model call, tool calls, validation |
| Correlation IDs | One ID threaded through app → model → tools → downstream systems |
| Token/cost telemetry | Input/output/thinking tokens and cost per request, per segment, per model |
| Retrieval telemetry | Query, retrieved chunk IDs, scores, rerank order |
| Errors & stop reasons | `stop_reason`, error codes, retries, fallbacks, degraded-mode flags |

```text
[req id: 9f3a] ── app ──▶ retrieve(k=50) ──▶ rerank(6) ──▶ opus-5 ──▶ tool:lookup ──▶ validate ──▶ resp
                 correlation id 9f3a threads through every span; token+cost tagged per span
```

Per-segment cost/latency telemetry is what lets you route, cache and optimise (D4) with evidence rather than guesses.

---

## 3.12 Enterprise integration patterns and data freshness

| Pattern | Purpose | Claude-specific note |
| --- | --- | --- |
| Queue | Decouple bursty producers from rate-limited model calls | Smooths load against RPM/ITPM tiers |
| Webhook | React to external events (ticket created, doc updated) | Trigger ingestion/refresh and agent runs |
| Idempotency | Safe retries without duplicate side effects | Idempotency keys on tool actions; critical with backoff retries |
| Batch | Latency-tolerant bulk work | **Message Batches API**: 50% discount, results within 24 h |
| Refresh / freshness | Keep the index current | Event- or schedule-driven re-chunk/re-embed/re-index (ties to 3.7) |

:::tip[Exam signal]
"Retried request charged/acted twice" → **idempotency keys**. "Thousands of documents to classify overnight, cost matters" → **Batch API**. "Document changed but the answer didn't" → **freshness/refresh pipeline** and re-indexing.
:::

---

## 3.13 A concrete RAG pipeline configuration

Item writers reward candidates who can read a config and predict its behaviour. A defensible enterprise baseline:

```yaml
ingestion:
  sources: [confluence, s3_pdfs, postgres_kb]
  dedup: content_hash
  pii_scrub: true
chunking:
  strategy: structural           # split on headings/sections
  fallback: recursive
  target_tokens: 400
  overlap_tokens: 60
  parent_child: true             # match child (~400), return parent (~1500)
embedding:
  model: text-embedding-3-large  # keep query + index model identical
  dimensions: 1024
index:
  store: managed_vector_db
  metadata: [tenant_id, doc_type, acl, updated_at, source_id]
  filters_before_search: [tenant_id, acl]   # security + precision
retrieval:
  mode: hybrid                   # BM25 + dense, RRF fusion
  k: 50
  rerank:
    model: cross_encoder_reranker
    keep_top: 6
  query_transform: [rewrite, multi_query]   # for conversational/sparse queries
generation:
  model: claude-sonnet-5
  grounding: answer_only_from_context
  citations: true
  on_insufficient_context: "say so; do not use parametric memory"
refresh:
  trigger: [webhook_on_update, nightly_schedule]
  action: re_chunk_re_embed_re_index
```

| Knob | Effect if too low | Effect if too high |
| --- | --- | --- |
| `target_tokens` | fragments ideas; loses context | dilutes relevance; buries the answer |
| `overlap_tokens` | boundary facts lost | storage/cost bloat, duplicate hits |
| `k` (pre-rerank) | low recall | noise; slower rerank |
| `keep_top` | best passage may be dropped | context bloat, higher cost, needle dilution |

:::tip[Exam signal]
"Retrieve wide (k≈50 hybrid) → rerank → keep 6 → ground with citations" is the canonical high-precision recipe. Distractors that skip reranking, use dense-only, or raise `keep_top` to 50 are the wrong answers.
:::

---

## 3.14 Retrieval evaluation with worked numbers

You must be able to *compute* the metrics, not just name them.

**Setup.** 5 queries; for each we know the single relevant chunk and where it ranked in the retrieved list:

| Query | Rank of relevant chunk | In top-3? | Reciprocal rank |
| --- | --- | --- | --- |
| Q1 | 1 | yes | 1/1 = 1.00 |
| Q2 | 4 | no | 1/4 = 0.25 |
| Q3 | 2 | yes | 1/2 = 0.50 |
| Q4 | (not retrieved) | no | 0 |
| Q5 | 1 | yes | 1/1 = 1.00 |

```text
recall@3 = (#queries whose relevant chunk is in top-3) / total = 3/5 = 0.60
MRR      = mean reciprocal rank = (1.00 + 0.25 + 0.50 + 0 + 1.00) / 5 = 0.55
```

Now suppose adding a **reranker** moves Q2's relevant chunk to rank 2 and retrieves Q4's at rank 3:

```text
recall@3 → 5/5 = 1.00   (both now in top-3)
MRR      → (1.00 + 0.50 + 0.50 + 0.333 + 1.00)/5 = 0.667
```

The reranker lifted recall@3 from 0.60 → 1.00 and MRR from 0.55 → 0.67 — the single biggest precision/ranking win, without touching chunking or embeddings.

| Metric | Formula | Reads as |
| --- | --- | --- |
| recall@k | relevant-in-top-k ÷ total queries | "did we retrieve the answer at all?" |
| MRR | mean(1 ÷ rank of first relevant) | "how high did it rank?" |
| precision@k | relevant-in-top-k ÷ k | "how clean is the top-k?" |
| faithfulness | supported claims ÷ total claims | "did the answer stay grounded?" |

:::caution[Interpreting the numbers]
High **recall** but low **faithfulness** ⇒ retrieval is fine, generation/grounding is broken. Low **recall** ⇒ fix chunking/embeddings/retrieval first. Reporting only end-to-end accuracy hides which of these is true — the aggregate-metric trap.
:::

---

## 3.15 Multi-tenant isolation and ACL-aware retrieval

In a shared index, retrieval is a **security boundary**. A query must never surface another tenant's or an unauthorised user's chunk.

```text
user (tenant_42, role: support) asks a question
   │  attach identity: {tenant_id: 42, acl_groups: [support]}
   ▼
filter BEFORE/ALONGSIDE vector search:
   WHERE tenant_id = 42 AND acl IN ('public','support')
   ▼
hybrid retrieve → rerank → ground   (only authorised chunks ever enter context)
```

| Gap | Leak | Control |
| --- | --- | --- |
| No `tenant_id` filter | Cross-tenant data exposure | Mandatory pre-filter on `tenant_id` |
| ACL applied only in the prompt | Model can be talked past it | Enforce ACL at the **retrieval layer**, not in prose |
| Shared service account for retrieval | User sees data they can't access | Propagate end-user identity into the query filter |
| Rerank/cache ignores tenant | Cached cross-tenant hit | Key caches by tenant + ACL |

:::danger[Retrieval-time filtering is the control]
Filtering *after* generation ("the model won't mention it") is not a control — the unauthorised chunk already entered context and can leak. The exam's correct answer filters by `tenant_id`/ACL **before** the vector search.
:::

---

## 3.16 Scenario walkthrough: a leaking, stale, over-privileged support RAG agent

**Scenario.** A B2B SaaS support agent serves 300 tenants from one shared vector index. Three complaints arrive: (1) a tenant occasionally sees another tenant's runbook in an answer; (2) after customers update their docs, the agent still quotes yesterday's version, confidently; (3) the agent has 22 tools including `delete_ticket` and `refund` it should never call. Retrieval is dense-only, k=8, no reranker.

**Expert reasoning trace.**

<Steps>

1. **Isolation first (highest severity).** Cross-tenant exposure is a security incident. Add a **mandatory `tenant_id` (and ACL) pre-filter** on every query, key caches by tenant, and propagate end-user identity. Do not rely on prompt wording to keep tenants apart.

2. **Freshness second.** Confident-wrong-after-update is stale retrieval: the refresh updated the source store but not the **vector index**. Inspect retrieved chunk IDs (they'll be old), then fix the **re-chunk/re-embed/re-index** pipeline with webhook + nightly triggers.

3. **Least privilege third.** Remove `delete_ticket`, `refund`, and other unneeded tools from the allowlist — do not merely log or add "are you sure?". Cut to ~4–5 tools; use tool search + `defer_loading` if the legitimate catalogue is large.

4. **Then quality.** Dense-only + k=8 + no rerank underperforms on exact identifiers and precision. Move to **hybrid retrieval, k=50, rerank, keep top 6**, and measure recall@k and faithfulness per tenant segment.

5. **Close the loop.** Instrument correlation IDs and per-segment retrieval telemetry so the next incident is reconstructable.

</Steps>

**Why the tempting alternatives are wrong:** "add a system-prompt rule not to reveal other tenants" is prompt-as-enforcement and leaves the leak; "switch to a bigger model" fixes neither isolation nor freshness; "log the dangerous tools" leaves excessive agency; "raise k to 500" adds noise instead of adding a reranker.

---

## 3.17 Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "Confident-but-wrong means the model is bad." | After a data change it usually means stale retrieval/indexing. | Inspect retrieval first; prompt/model fixes are distractors. |
| "Dense embeddings retrieve everything." | Dense is weak on exact IDs/codes; hybrid adds sparse. | Part-number/SKU stems require hybrid. |
| "A bigger context window replaces RAG." | It costs more per call, can't cite, and degrades on the needle. | Long-context-stuffing is the wrong answer for large/changing corpora. |
| "Fine-tuning is how you add knowledge." | Fine-tuning bakes behaviour/style; changing facts need RAG. | Weekly-changing-facts stems reject fine-tuning. |
| "Logging a dangerous tool makes it safe." | The capability is still reachable — excessive agency. | Least privilege means *removal*, not observation. |
| "Prompt rules keep tenants isolated." | Isolation must be enforced at the retrieval filter. | Retrieval-time ACL filtering is the correct control. |
| "MCP servers are authenticated by default." | Remote MCP needs OAuth 2.1 + per-user checks. | Unauthenticated remote MCP is a security trap. |
| "Retrying is always safe." | Non-idempotent actions double up; add idempotency keys. | Double-charge stems test idempotency. |

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| Blaming the prompt/model for confident-wrong answers after a refresh | The cause is usually stale retrieval/indexing; inspect what was retrieved first |
| Fixing an unneeded dangerous tool by logging or confirming it | Leaves excessive agency; least privilege means **removing** the tool |
| Giving an agent 18 tools "for flexibility" | Slower, error-prone; cut to 4–5, use tool search + `defer_loading` |
| Using dense-only retrieval for part numbers / exact IDs | Dense misses exact tokens; use sparse/hybrid |
| Measuring only end-to-end accuracy | Hides whether retrieval or generation failed; measure recall@k and faithfulness separately |
| Long-context stuffing a large, changing corpus | Costly, no citations, needle degradation; use RAG |
| Fine-tuning to inject frequently-changing facts | Retraining cadence infeasible; use RAG |
| Skipping reranking | Best passage buried below top-k; precision suffers |
| One service account for all users' tool calls | Cross-user data exposure; propagate identity + per-user ACLs |
| Unauthenticated remote MCP server | Anyone can invoke powerful tools; require OAuth 2.1 |
| Retrying non-idempotent tool actions without keys | Duplicate side effects (double refund/charge) |
| Loading all tools/docs into context up front | Expensive, cache-hostile, worse tool selection; use progressive discovery |
| Enforcing multi-tenant isolation with a system-prompt rule | Isolation must be a retrieval-time `tenant_id`/ACL filter, not prose |
| Filtering unauthorised chunks after generation | The chunk already entered context; filter before the vector search |
| Raising k to hundreds instead of adding a reranker | Adds noise; reranking is the precision/ranking win |
| Reporting only end-to-end accuracy for a RAG system | Hides whether retrieval or grounding failed; measure recall@k and faithfulness |
| Using the same embedding model for query and index inconsistently | Query/index model mismatch wrecks similarity; keep them identical |
| Caching retrieval results without keying by tenant/ACL | Risks serving a cross-tenant cached hit |

---

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · A policy assistant kept answering with the OLD figure after a policy document was updated last night, and it sounds completely confident. What should the architect investigate FIRST? (Select one)">
    A. Rewrite the system prompt to be more accurate.
    B. Inspect what retrieval actually returned; the updated document was likely not re-chunked/re-embedded/re-indexed, so retrieval is serving stale vectors.
    C. Switch to Opus 5 with xhigh effort.
    D. Add more few-shot examples.

    **Answer: B.** Confident-but-wrong immediately after a refresh points to stale retrieval/indexing, not the model. The first move is to log the retrieved chunks and confirm whether the refresh updated the vector index. Prompt rewrites (A, D) and a bigger model (C) cannot fix stale retrieval.
  </AccordionItem>

  <AccordionItem title="Q2 · A support agent has 18 tools, including `delete_account` and `issue_refund`, which its role should never use. What is the correct remediation? (Select one)">
    A. Keep the tools but log every call for audit.
    B. Remove the unneeded tools from the agent's allowlist (least privilege) so it cannot invoke them at all.
    C. Add a confirmation prompt before those tools run.
    D. Add a system-prompt sentence forbidding their use.

    **Answer: B.** Least privilege means the capability should not be present. Removing the tools eliminates the excessive agency. Logging (A) and confirmation (C) leave the capability reachable; a prompt rule (D) is prompt-as-enforcement and can be bypassed.
  </AccordionItem>

  <AccordionItem title="Q3 · Users search a parts catalogue by exact part numbers AND by descriptions. Dense-only retrieval misses many exact-number queries. What is the BEST fix? (Select one)">
    A. Increase k to 500.
    B. Use hybrid retrieval (BM25 + dense with rank fusion) so exact identifiers and semantic matches both rank well.
    C. Switch to a bigger generation model.
    D. Remove metadata filters.

    **Answer: B.** Dense embeddings are weak on exact tokens/codes; sparse (BM25) handles them, and hybrid fusion covers both query types. A huge k (A) adds noise without fixing lexical matching; a bigger model (C) doesn't change what is retrieved; removing filters (D) hurts precision and security.
  </AccordionItem>

  <AccordionItem title="Q4 · Retrieval eval shows recall@10 = 0.95 but faithfulness is low and answers include facts not in the retrieved chunks. Where is the failure and the fix? (Select one)">
    A. Retrieval; lower k.
    B. Generation/grounding; strengthen the instruction to answer only from context, add citations, and consider reranking so the best passage is on top.
    C. Embeddings; change the model.
    D. Indexing; re-index everything.

    **Answer: B.** High recall means the right chunks are retrieved, so the fault is in generation/grounding — the model is answering from parametric memory. Grounding instructions, citations and reranking address it. The retrieval-side fixes (A, C, D) target a stage that is already performing well.
  </AccordionItem>

  <AccordionItem title="Q5 · A 2M-document knowledge base changes weekly and answers must cite the exact source clause. Which approach is BEST? (Select one)">
    A. Fine-tune a model on the corpus each week.
    B. RAG with hybrid retrieval, reranking and citations, plus a scheduled refresh pipeline.
    C. Stuff the whole corpus into the 1M context per query.
    D. Long context plus fine-tuning combined.

    **Answer: B.** Large, frequently-changing, citation-requiring corpora are the canonical RAG case. Weekly fine-tuning (A) is an infeasible retraining cadence for facts. The corpus exceeds the window and would cost too much and lose citations (C). (D) inherits both problems.
  </AccordionItem>

  <AccordionItem title="Q6 · A contract-QA system chunks contracts every 500 tokens with fixed size. Answers cite the wrong sub-clause and lose surrounding context. Which TWO changes help most? (Select two)">
    A. Use structural/document-aware chunking that splits on clauses/sections.
    B. Use parent–child chunking: match on small child chunks but return the larger parent clause for context.
    C. Switch to dense-only retrieval.
    D. Increase temperature.
    E. Remove citations.

    **Answer: A and B.** Contracts are structured, so clause/section-aware chunking preserves boundaries, and parent–child gives precise matching with enough surrounding context. Dense-only (C), temperature (D) and removing citations (E) do not address the chunking problem — the last two make it worse.
  </AccordionItem>

  <AccordionItem title="Q7 · A remote MCP server exposes powerful tools over Streamable HTTP with no auth. What must the architect add? (Select one)">
    A. Nothing; MCP is safe by default.
    B. OAuth 2.1 authentication on the remote MCP server, plus per-user permission checks inside the tools.
    C. A longer system prompt.
    D. A higher rate-limit tier.

    **Answer: B.** Remote MCP servers require OAuth 2.1, and tools must enforce per-user permissions so the agent acts with the caller's authority. MCP is not authenticated by default (A); prompts (C) and rate limits (D) do not address authorization.
  </AccordionItem>

  <AccordionItem title="Q8 · A nightly job must classify 200,000 documents; latency is not important but cost is. Which mechanism is BEST? (Select one)">
    A. Real-time synchronous calls in a tight loop.
    B. The Message Batches API for a 50% discount with results within 24 hours.
    C. A multi-agent system.
    D. Fine-tuning.

    **Answer: B.** Latency-tolerant bulk work is exactly the Batch API's use case (50% discount, results within 24 h). Synchronous loops (A) hit rate limits and cost more. Multi-agent (C) adds cost/complexity; fine-tuning (D) is unrelated to a classification batch.
  </AccordionItem>

  <AccordionItem title="Q9 · After enabling backoff retries, some refunds are issued twice. What is the correct fix? (Select one)">
    A. Disable retries entirely.
    B. Add idempotency keys to the refund tool so retried calls are de-duplicated and produce no duplicate side effect.
    C. Lower the model temperature.
    D. Log the duplicates and reconcile later.

    **Answer: B.** Retries on non-idempotent actions cause duplicate side effects; idempotency keys make retries safe. Disabling retries (A) harms resilience. Temperature (C) is irrelevant. Reconciling after the fact (D) still charged customers twice.
  </AccordionItem>

  <AccordionItem title="Q10 · A single retrieval sometimes isn't enough — some questions need follow-up lookups combining multiple sources. Which design fits, and what is the trade-off? (Select one)">
    A. Agentic RAG, where the model decides when/what to retrieve and can issue follow-up queries, at the cost of more latency and tokens.
    B. Stuff everything into context to avoid retrieval.
    C. Fine-tune on the multi-hop questions.
    D. Remove reranking to speed things up.

    **Answer: A.** Multi-hop, exploratory queries justify agentic RAG's model-driven retrieval loop; the trade-off is higher cost and latency versus one-shot RAG. Context stuffing (B) doesn't scale; fine-tuning (C) can't hold changing facts; removing reranking (D) hurts precision.
  </AccordionItem>

  <AccordionItem title="Q11 · A capability will be reused by many agents and clients across the company and must follow a standard protocol. Which integration mechanism is BEST? (Select one)">
    A. Hard-code it as a per-app CLI tool in each service.
    B. Build it as an MCP server (with OAuth 2.1 if remote) so many clients reuse it via a standard protocol.
    C. Implement it as an agent-to-agent handoff only.
    D. Paste its logic into every system prompt.

    **Answer: B.** Reuse across many clients under a standard protocol is MCP's purpose. Per-app CLI tools (A) fragment the implementation; agent-to-agent (C) is for specialised context isolation, not shared capability; prompt-pasting (D) is unmaintainable.
  </AccordionItem>

  <AccordionItem title="Q12 · An agent loads all 40 tools and 30 reference documents into context on every request; cost is high, cache hit rate is low, and tool selection is error-prone. What is the BEST remedy? (Select two)">
    A. Use tool search with `defer_loading: true` so only relevant tools are loaded.
    B. Package reference material as Skills loaded progressively on demand.
    C. Increase the context window to 1M and keep loading everything.
    D. Escalate every request to Opus 5.
    E. Disable prompt caching.

    **Answer: A and B.** Progressive discovery — tool search with deferred loading and on-demand Skills — keeps context lean, restores a stable cache prefix, and improves tool selection. A bigger window (C) still pays for the bloat; escalating models (D) raises cost; disabling caching (E) is the opposite of the fix.
  </AccordionItem>

  <AccordionItem title="Q13 · A B2B assistant serves 300 tenants from one shared vector index; a tenant occasionally sees another tenant's document in an answer. What is the correct control? (Select one)">
    A. Add a system-prompt rule telling the model not to reveal other tenants' data.
    B. Apply a mandatory `tenant_id` (and ACL) filter at retrieval time, before/alongside the vector search, so only authorised chunks ever enter context.
    C. Filter the answer after generation to remove other tenants' data.
    D. Give each tenant a bigger model.

    **Answer: B.** Isolation is a retrieval-time security boundary: filter by `tenant_id`/ACL before the vector search. A prompt rule (A) is bypassable prompt-as-enforcement; post-generation filtering (C) is too late — the chunk already entered context; a bigger model (D) doesn't isolate data.
  </AccordionItem>

  <AccordionItem title="Q14 · Across 5 queries the relevant chunk ranked 1, 4, 2, not-retrieved, 1. What are recall@3 and MRR? (Select one)">
    A. recall@3 = 1.00; MRR = 1.00.
    B. recall@3 = 0.60; MRR = 0.55.
    C. recall@3 = 0.55; MRR = 0.60.
    D. recall@3 = 0.80; MRR = 0.70.

    **Answer: B.** Three of five relevant chunks are in the top 3 (ranks 1, 2, 1) → recall@3 = 3/5 = 0.60. Reciprocal ranks are 1, 0.25, 0.5, 0, 1 → MRR = 2.75/5 = 0.55. Option A ignores the misses; C swaps the two values; D is arithmetically wrong.
  </AccordionItem>

  <AccordionItem title="Q15 · recall@10 = 0.62 (low) and answers are frequently missing the needed fact. Where should the architect work FIRST? (Select one)">
    A. Grounding; tighten the answer-only-from-context instruction.
    B. Retrieval: fix chunking/embeddings/hybrid and add reranking, because low recall means the relevant chunk often isn't retrieved at all.
    C. Add more citations.
    D. Switch the generation model to Opus 5.

    **Answer: B.** Low recall means the right chunk isn't reaching the top-k, so the failure is upstream in retrieval. Grounding/citations (A, C) and a bigger generation model (D) can't help if the answer was never retrieved.
  </AccordionItem>

  <AccordionItem title="Q16 · A RAG system indexes with `text-embedding-3-large` but a new service queries with a different embedding model. Similarity scores look random. What is the cause and fix? (Select one)">
    A. The vector store is corrupt; rebuild hardware.
    B. Query/index embedding-model mismatch; use the identical embedding model for both indexing and querying.
    C. k is too low; raise it to 1000.
    D. The generation model is too small.

    **Answer: B.** Embeddings from different models live in different vector spaces, so cross-model similarity is meaningless; query and index must use the same embedding model. It isn't hardware (A); raising k (C) can't fix incompatible vectors; the generation model (D) is unrelated to retrieval similarity.
  </AccordionItem>

  <AccordionItem title="Q17 · An agent retrieves k=8 with dense-only and no reranker; precision is poor and exact part numbers are missed. Which TWO changes give the biggest quality lift? (Select two)">
    A. Switch to hybrid retrieval (BM25 + dense) so exact identifiers rank.
    B. Retrieve wide (k≈50) and add a reranker, keeping the top 6.
    C. Increase temperature.
    D. Remove citations to speed responses.
    E. Move to a 1M-token context and stuff everything.

    **Answer: A and B.** Hybrid retrieval fixes exact-identifier misses and retrieve-wide-then-rerank fixes precision — the two canonical levers. Temperature (C) is irrelevant to retrieval; removing citations (D) harms grounding traceability; stuffing context (E) inflates cost without improving ranking.
  </AccordionItem>

  <AccordionItem title="Q18 · Which TWO signals let you reconstruct a failed multi-step RAG request end to end? (Select two)">
    A. A correlation ID threaded through app → retrieval → model → tools → downstream.
    B. Per-span traces capturing retrieved chunk IDs, rerank order, token/cost, and `stop_reason`.
    C. Only the final HTTP status code.
    D. The model's own assessment that it did fine.
    E. A daily aggregate request count.

    **Answer: A and B.** A correlation ID plus per-span traces (with retrieval detail and cost) make an incident reconstructable. A status code (C) and a daily count (E) are too coarse; self-assessment (D) is unreliable (self-report anti-pattern).
  </AccordionItem>
</Accordions>

## Key takeaways

- RAG is a fixed pipeline: **ingest → chunk → embed → index → retrieve → rerank → ground → cite**; every decision slots into a stage.
- Match **chunking to data shape**: document-aware for structured text, parent–child/late for cross-referential prose, semantic for shifting topics.
- Use **hybrid** retrieval (dense + sparse) for real corpora; **rerank** for precision; add rewriting/HyDE/multi-query for recall.
- **Ground** answers to retrieved context, cite sources, and prefer "insufficient context" over fabrication.
- Evaluate retrieval and generation **separately** (recall@k, MRR vs faithfulness); confident-wrong-after-refresh means **inspect retrieval/indexing first**.
- Choose **RAG for changing/cited facts, long context for small stable corpora, fine-tuning for behaviour/style**.
- Apply **least privilege by removing** unneeded (especially destructive) tools; keep agents to ~4–5 tools with tool search for larger catalogues.
- Propagate **user identity** to tools, enforce per-user ACLs, and require **OAuth 2.1** on remote MCP servers.
- Instrument **traces, correlation IDs and token/cost telemetry**; use queues, idempotency keys, webhooks, Batch API and a **freshness pipeline** for enterprise integration.
- **Compute** retrieval metrics: recall@k = relevant-in-top-k ÷ queries; MRR = mean(1÷rank); a reranker typically lifts both the most.
- In a shared index, retrieval is a **security boundary**: filter by `tenant_id`/ACL **before** the vector search, propagate end-user identity, and key caches by tenant.
- Keep the **query and index embedding models identical**; the canonical recipe is retrieve wide (k≈50 hybrid) → rerank → keep ~6 → ground with citations.
