Appendix · OpenAI
RAG Cookbook (OpenAI)
Retrieval-augmented generation recipes for the OpenAI stack — pipeline anatomy, chunking trade-offs, embeddings and vector stores, the file search tool versus a self-managed store, hybrid retrieval and reranking, grounding, permissions, freshness, evaluating retrieval separately from generation, a failure taxonomy, and a cost and latency budget.
RAG grounds answers in retrieved passages at query time. Reach for it when the corpus is large or changing and you need citations. Prefer long context for a small stable corpus that fits comfortably in the window; prefer fine-tuning for style and format, not for facts. This cookbook mirrors the objectives of the Academy Build with Retrieval-Augmented Generation course; it is independent preparation built from published learning objectives.
Assessment signal
“Confident but wrong after a document refresh” points at retrieval or indexing first — a stale index, a changed chunker, re-embedded with a different model — not the prompt or the model. Rewriting the prompt is the distractor.
Pipeline anatomy
ingest ──► chunk ──► embed ──► store (build time, offline) │ query ──► rewrite ──► retrieve ──► rerank ──► assemble ──► generate ──► cite │ (query time, online)Two halves fail for different reasons. The build-time half (chunk, embed, store) fails silently: a bad chunker or a changed embedding model degrades every future query. The query-time half (retrieve, rerank, generate) fails visibly, per request. Instrument both, and evaluate them separately.
| Stage | Owns | Common failure |
|---|---|---|
| Chunk | Splitting docs into retrievable units | Splits mid-idea; drift after re-ingest |
| Embed | Turning text into vectors | Wrong or mismatched model between index and query |
| Store | Indexing and nearest-neighbour search | Stale vectors; missing permission metadata |
| Retrieve | Fetching candidates | Misses exact IDs (dense only); wrong top-k |
| Rerank | Ordering candidates by relevance | Absent, so the gold passage sits too low |
| Generate | Writing the grounded answer | Ignores context; no citation |
When RAG vs alternatives
| Situation | Choose | Why |
|---|---|---|
| Thousands of docs, updated often, must cite | RAG | Fresh, grounded, auditable |
| A 50-page handbook that rarely changes and fits the window | Long context | No retrieval infrastructure; simplest |
| “Always answer in our house style / format” | Fine-tuning or few-shot | Behaviour, not facts |
| Model must decide when and what to retrieve | Agentic RAG (retrieval as a tool) | Iterative reformulation |
The GPT-5.6 and GPT-6 models carry a large context window, which tempts teams to “just paste everything”. That works for a small stable corpus, but for a large or changing one it is more expensive per query, harder to keep fresh, and gives you no citations to audit.
Chunking strategies with trade-offs
| Strategy | Start here | Best for | Trade-off |
|---|---|---|---|
| Fixed-size | 512 tokens, 50-token overlap | Uniform prose | Splits mid-idea |
| Recursive (by separator) | 400–800 tokens, split on \n\n then \n then sentence | Mixed prose | Uneven sizes |
| Semantic | Break where embedding similarity dips | Topic-shifting docs | Compute cost at ingest |
| Structural / document-aware | One chunk per section or heading | Manuals, contracts, wikis | Needs clean structure |
| Parent-child (small-to-big) | Child 150–300 tok to match, return parent 800–1200 tok | Precise hit plus broad context | Two-tier store |
Overlap rule of thumb: 10–20% of the chunk size. Too little loses boundary context; too much inflates the index and produces duplicate hits.
Chunk drift
When documents are re-ingested with a different chunker or size, retrieval quality shifts silently and answers go confidently wrong. Version your chunking configuration and re-run retrieval evals after any change to it.
Embeddings and vector stores
Embeddings turn text into vectors so that semantically similar passages sit near each other. Two rules dominate:
- The index and the query must use the same embedding model and dimension. Mixing models is the classic silent corruption: cosine similarity between vectors from two different models is meaningless.
- Re-embed the whole corpus when you change the model. A half-migrated index returns garbage for the un-migrated half. Treat an embedding-model change like a schema migration: all or nothing, gated on an eval.
The store holds vectors plus metadata (source, section, ACL, timestamp) and does approximate nearest-neighbour search. On the OpenAI stack you can either let the platform manage this for you or run your own.
The file search tool vs a self-managed store
| File search tool (managed) | Self-managed vector store | |
|---|---|---|
| Who chunks and embeds | The platform | You |
| Who indexes and retrieves | The platform, called as a tool | Your database and code |
| Control over chunking / reranking | Limited to the tool’s options | Full |
| Best when | You want retrieval fast, with less infrastructure | You need custom chunking, hybrid search, your own reranker, or your own store of record |
| Trade-off | Less tuning surface | You own freshness, permissions and evals end to end |
The decision rule: reach for the file search tool when you want grounded answers over your files without standing up a retrieval stack, and reach for a self-managed store when you need control the tool does not expose — a specific chunking scheme, hybrid retrieval, a cross-encoder reranker, or permission-aware retrieval integrated with your identity system.
Hybrid retrieval and reranking
Dense (embedding) retrieval catches paraphrase; sparse (BM25) retrieval catches exact IDs, error codes and rare terms. Combine them, then rerank.
retrieval: dense: { top_k: 40 } sparse: { algorithm: bm25, top_k: 40 } fusion: { method: reciprocal_rank_fusion, k: 60 } rerank: { model: cross-encoder, top_n: 8 }Reciprocal Rank Fusion: a document’s score is the sum over lists of 1 / (k + rank). With k=60, a document ranked #1 in both lists scores about 1/61 + 1/61 = 0.033; ranked #1 in one list and absent in the other scores about 1/61 = 0.016. RRF needs no score calibration between the two systems — it fuses by rank position alone.
A cross-encoder reranker reads the query and each candidate together, giving far better precision than the first-stage retriever, at higher per-query cost. Retrieve broadly (top-k 40–100), then rerank down to a small top-n (5–10) for the model.
| First-stage retrieval | Cross-encoder rerank | |
|---|---|---|
| Encodes | Query and docs separately, precomputed | Query and doc jointly, per pair |
| Speed | Fast, index-time | Slow, query-time |
| Strength | Recall | Precision / ordering |
| Role | Fetch top-k | Order to top-n |
Grounding and citation formatting
Grounding is only real if the answer can be traced to a passage. Two disciplines:
- Constrain the answer to the retrieved context, and provide a sentinel for “not found”:
Answer only from the passages below. If they do not contain the answer, reply NOT COVERED.The sentinel lets code detect a non-answer instead of shipping a guess. - Return citations the reader can verify: a source identifier and a section or span for each claim. When the model must synthesise across passages, ask it to cite each supporting passage rather than one blanket citation at the end.
Permission-aware retrieval
The retriever must never surface a passage the asking user is not allowed to see. This is an ingest-and-retrieve concern, not a prompt concern.
- Attach access metadata at ingest — store each chunk’s ACL (owner, group, sensitivity) alongside its vector.
- Filter at query time by the asking user’s identity — apply the permission filter as part of retrieval, before the model ever sees a candidate.
- Never rely on a prompt instruction like “do not reveal restricted content” — an unfiltered restricted passage in context is already a leak, whatever the prompt says.
- Re-check on permission changes — when a user loses access, their future queries must stop returning those chunks; permission state lives with the data, not the session.
Permissions are retrieval, not prompting
The trap is a stem where sensitive data leaks and the tempting fix is a stronger system-prompt rule. If a restricted passage was retrieved into context, the prompt is irrelevant — the fix is a permission filter at retrieval time.
Freshness
A RAG system is only as current as its index. Decide, per corpus, how fresh it must be and build for that:
| Freshness need | Approach |
|---|---|
| Minutes (prices, tickets) | Retrieve from the source of record at query time, or stream updates into the index |
| Hours to a day | Scheduled re-ingest; version the index and swap atomically |
| Rarely changes | Periodic full rebuild; gate the swap on a retrieval eval |
Whatever the cadence, gate re-ingests on an eval so a chunker or model change cannot silently degrade retrieval, and store a build timestamp so “confident but wrong after a refresh” is diagnosable.
Evaluating retrieval separately from generation
A good answer can hide bad retrieval and vice versa, so measure them apart.
| Metric | Measures | Example |
|---|---|---|
| Recall@k | Was the gold passage in the top-k? | 8 of 10 queries had it in top-5 → 0.80 |
| MRR | Rank of the first relevant hit | Ranks 1, 3, 2 → (1 + 1/3 + 1/2)/3 = 0.61 |
| Precision@k | Fraction of top-k that are relevant | 2 of top-5 relevant → 0.40 |
| Faithfulness | Is the answer supported by the retrieved context? | 47 of 50 claims grounded → 0.94 |
| Answer relevance | Does the answer address the question? | Rubric or model-graded |
Worked comparison on a 200-query set:
| Config | Recall@5 | MRR | Faithfulness | Answer relevance |
|---|---|---|---|---|
| Dense only | 0.71 | 0.52 | 0.86 | 0.83 |
| Hybrid (RRF) | 0.83 | 0.61 | 0.90 | 0.86 |
| Hybrid + reranker | 0.83 | 0.74 | 0.94 | 0.90 |
Reranking barely moves Recall@5 (same candidates) but lifts MRR and faithfulness sharply by putting the right passage first, where the model reads it earliest. Report every metric per segment (document type, source, language) — an aggregate 0.90 can hide one source at 0.40.
Failure taxonomy with fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Answer wrong, gold passage not retrieved | Chunk drift, wrong embedding model, poor query | Version chunking; align embed models; add query rewriting |
| Gold passage retrieved but ranked low | No reranker; bad fusion weights | Add a cross-encoder reranker; tune RRF k |
| Exact codes / IDs missed | Dense-only retrieval | Add BM25 (hybrid) |
| Retrieved and top-ranked but answer wrong | Model ignoring context; context not passed | Check the prompt actually includes chunks; constrain to context |
| Confident but wrong after a refresh | Stale or re-chunked index | Re-run retrieval evals; check build timestamp |
| One source consistently bad | Aggregate hid it | Report per-segment; fix that source’s ingest |
| Restricted content surfaced | No permission filter at retrieval | Filter by identity at query time |
Cost and latency budget worked example
A support assistant over 40,000 articles, target end-to-end latency under 2 seconds at p95, gpt-5.6-terra for generation.
Per query budget (p95 target: 1900 ms) query rewrite (gpt-5.6-luna, low effort) ~120 ms hybrid retrieve (dense + BM25, top_k 40) ~140 ms cross-encoder rerank (40 -> 8) ~180 ms generate grounded answer (terra, 8 chunks) ~1200 ms citation assembly (code) ~20 ms -------- total (p95) ~1660 ms ✓ under 1900Cost levers, cheapest first: cut the number of chunks fed to the model (8 → 5 with a stronger reranker) before touching the model tier; use gpt-5.6-luna for the rewrite step; cache the stable system prefix on the API; and reserve a larger model for only the queries a router flags as hard. The wrong instinct — “use a bigger model to fix wrong answers” — spends money on a generation problem that is usually a retrieval problem.
Common misconceptions
| Misconception | Reality | Why it matters on the assessment |
|---|---|---|
| “More context always helps” | Noise lowers faithfulness; rerank to a small top-n | Over-retrieval distractor |
| “Dense retrieval covers everything” | It misses exact IDs and codes; add BM25 | Hybrid-search signal |
| “Wrong answer means fix the prompt” | Check retrieval first after a refresh | Root-cause distractor |
| “A prompt rule keeps restricted data out” | Permissions are enforced at retrieval, not by prompt | Permission trap |
| “Fine-tune to add new facts” | Fine-tuning is for style; RAG for facts | RAG-vs-fine-tuning distractor |
| “Reranking improves recall” | It improves ordering and precision, not recall | Metric-confusion distractor |
Scenario walkthrough
A team runs a KB assistant over 40,000 articles refreshed weekly. This week it began giving confident wrong answers. Users ask both natural-language questions and exact error codes. What is the FIRST step and the right architecture?
- FIRST step — pull traces and check whether the gold passage is even retrieved. It surfaced that this week’s re-ingest changed the chunker (drift). This is a retrieval problem, so rewriting the prompt would waste the cycle.
- Chunking — structural or parent-child for articles: child chunks for precise matches, parent sections for context; version the config and gate re-ingests on an eval.
- Retrieval — hybrid dense + BM25, because error codes need the exact-term matching that dense retrieval misses.
- Reranking — cross-encoder to top-8; the gold article was retrieved but sat at rank 14 before reranking.
- Permissions — filter by the asking user’s identity at query time, since some articles are internal-only.
- Eval — recall@k, MRR and faithfulness, reported per source, so a single failing source cannot hide behind the aggregate.
Rejected alternatives: rewriting the system prompt (retrieval was broken), fine-tuning on the KB (facts change weekly — that is RAG’s job), dense-only retrieval (misses error codes), and trusting the aggregate score (it masked the failing source).
Key takeaways
- RAG for large or changing corpora with citations; long context for small stable corpora; fine-tuning for style, not facts.
- Chunk to match the data shape; parent-child balances precise hits with context; version the config and gate re-ingests.
- Keep the index and query on the same embedding model; re-embed the whole corpus on any model change.
- Choose the file search tool for speed with less infrastructure, a self-managed store when you need chunking, hybrid retrieval, a reranker or permission integration you control.
- Hybrid (dense + BM25) plus a cross-encoder reranker beats either alone; RRF fuses without score calibration.
- Enforce permissions and freshness at retrieval time, and evaluate retrieval (recall@k, MRR) separately from generation (faithfulness), per segment.
Last updated Sep 18, 2026