AI Cert Prep
Type to search documentation.

Appendix · OpenAI

RAG Cookbook (OpenAI)

Retrieval-augmented generation recipes for the OpenAI stack — pipeline anatomy, chunking trade-offs, embeddings and vector stores, the file search tool versus a self-managed store, hybrid retrieval and reranking, grounding, permissions, freshness, evaluating retrieval separately from generation, a failure taxonomy, and a cost and latency budget.

RAG grounds answers in retrieved passages at query time. Reach for it when the corpus is large or changing and you need citations. Prefer long context for a small stable corpus that fits comfortably in the window; prefer fine-tuning for style and format, not for facts. This cookbook mirrors the objectives of the Academy Build with Retrieval-Augmented Generation course; it is independent preparation built from published learning objectives.

Assessment signal

“Confident but wrong after a document refresh” points at retrieval or indexing first — a stale index, a changed chunker, re-embedded with a different model — not the prompt or the model. Rewriting the prompt is the distractor.

Pipeline anatomy

text
ingest ──► chunk ──► embed ──► store (build time, offline)
│
query ──► rewrite ──► retrieve ──► rerank ──► assemble ──► generate ──► cite
│
(query time, online)

Two halves fail for different reasons. The build-time half (chunk, embed, store) fails silently: a bad chunker or a changed embedding model degrades every future query. The query-time half (retrieve, rerank, generate) fails visibly, per request. Instrument both, and evaluate them separately.

StageOwnsCommon failure
ChunkSplitting docs into retrievable unitsSplits mid-idea; drift after re-ingest
EmbedTurning text into vectorsWrong or mismatched model between index and query
StoreIndexing and nearest-neighbour searchStale vectors; missing permission metadata
RetrieveFetching candidatesMisses exact IDs (dense only); wrong top-k
RerankOrdering candidates by relevanceAbsent, so the gold passage sits too low
GenerateWriting the grounded answerIgnores context; no citation

When RAG vs alternatives

SituationChooseWhy
Thousands of docs, updated often, must citeRAGFresh, grounded, auditable
A 50-page handbook that rarely changes and fits the windowLong contextNo retrieval infrastructure; simplest
“Always answer in our house style / format”Fine-tuning or few-shotBehaviour, not facts
Model must decide when and what to retrieveAgentic RAG (retrieval as a tool)Iterative reformulation

The GPT-5.6 and GPT-6 models carry a large context window, which tempts teams to “just paste everything”. That works for a small stable corpus, but for a large or changing one it is more expensive per query, harder to keep fresh, and gives you no citations to audit.

Chunking strategies with trade-offs

StrategyStart hereBest forTrade-off
Fixed-size512 tokens, 50-token overlapUniform proseSplits mid-idea
Recursive (by separator)400–800 tokens, split on \n\n then \n then sentenceMixed proseUneven sizes
SemanticBreak where embedding similarity dipsTopic-shifting docsCompute cost at ingest
Structural / document-awareOne chunk per section or headingManuals, contracts, wikisNeeds clean structure
Parent-child (small-to-big)Child 150–300 tok to match, return parent 800–1200 tokPrecise hit plus broad contextTwo-tier store

Overlap rule of thumb: 10–20% of the chunk size. Too little loses boundary context; too much inflates the index and produces duplicate hits.

Chunk drift

When documents are re-ingested with a different chunker or size, retrieval quality shifts silently and answers go confidently wrong. Version your chunking configuration and re-run retrieval evals after any change to it.

Embeddings and vector stores

Embeddings turn text into vectors so that semantically similar passages sit near each other. Two rules dominate:

  1. The index and the query must use the same embedding model and dimension. Mixing models is the classic silent corruption: cosine similarity between vectors from two different models is meaningless.
  2. Re-embed the whole corpus when you change the model. A half-migrated index returns garbage for the un-migrated half. Treat an embedding-model change like a schema migration: all or nothing, gated on an eval.

The store holds vectors plus metadata (source, section, ACL, timestamp) and does approximate nearest-neighbour search. On the OpenAI stack you can either let the platform manage this for you or run your own.

The file search tool vs a self-managed store

File search tool (managed)Self-managed vector store
Who chunks and embedsThe platformYou
Who indexes and retrievesThe platform, called as a toolYour database and code
Control over chunking / rerankingLimited to the tool’s optionsFull
Best whenYou want retrieval fast, with less infrastructureYou need custom chunking, hybrid search, your own reranker, or your own store of record
Trade-offLess tuning surfaceYou own freshness, permissions and evals end to end

The decision rule: reach for the file search tool when you want grounded answers over your files without standing up a retrieval stack, and reach for a self-managed store when you need control the tool does not expose — a specific chunking scheme, hybrid retrieval, a cross-encoder reranker, or permission-aware retrieval integrated with your identity system.

Hybrid retrieval and reranking

Dense (embedding) retrieval catches paraphrase; sparse (BM25) retrieval catches exact IDs, error codes and rare terms. Combine them, then rerank.

yaml
retrieval:
dense: { top_k: 40 }
sparse: { algorithm: bm25, top_k: 40 }
fusion: { method: reciprocal_rank_fusion, k: 60 }
rerank: { model: cross-encoder, top_n: 8 }

Reciprocal Rank Fusion: a document’s score is the sum over lists of 1 / (k + rank). With k=60, a document ranked #1 in both lists scores about 1/61 + 1/61 = 0.033; ranked #1 in one list and absent in the other scores about 1/61 = 0.016. RRF needs no score calibration between the two systems — it fuses by rank position alone.

A cross-encoder reranker reads the query and each candidate together, giving far better precision than the first-stage retriever, at higher per-query cost. Retrieve broadly (top-k 40–100), then rerank down to a small top-n (5–10) for the model.

First-stage retrievalCross-encoder rerank
EncodesQuery and docs separately, precomputedQuery and doc jointly, per pair
SpeedFast, index-timeSlow, query-time
StrengthRecallPrecision / ordering
RoleFetch top-kOrder to top-n

Grounding and citation formatting

Grounding is only real if the answer can be traced to a passage. Two disciplines:

  • Constrain the answer to the retrieved context, and provide a sentinel for “not found”: Answer only from the passages below. If they do not contain the answer, reply NOT COVERED. The sentinel lets code detect a non-answer instead of shipping a guess.
  • Return citations the reader can verify: a source identifier and a section or span for each claim. When the model must synthesise across passages, ask it to cite each supporting passage rather than one blanket citation at the end.

Permission-aware retrieval

The retriever must never surface a passage the asking user is not allowed to see. This is an ingest-and-retrieve concern, not a prompt concern.

  1. Attach access metadata at ingest — store each chunk’s ACL (owner, group, sensitivity) alongside its vector.
  2. Filter at query time by the asking user’s identity — apply the permission filter as part of retrieval, before the model ever sees a candidate.
  3. Never rely on a prompt instruction like “do not reveal restricted content” — an unfiltered restricted passage in context is already a leak, whatever the prompt says.
  4. Re-check on permission changes — when a user loses access, their future queries must stop returning those chunks; permission state lives with the data, not the session.

Permissions are retrieval, not prompting

The trap is a stem where sensitive data leaks and the tempting fix is a stronger system-prompt rule. If a restricted passage was retrieved into context, the prompt is irrelevant — the fix is a permission filter at retrieval time.

Freshness

A RAG system is only as current as its index. Decide, per corpus, how fresh it must be and build for that:

Freshness needApproach
Minutes (prices, tickets)Retrieve from the source of record at query time, or stream updates into the index
Hours to a dayScheduled re-ingest; version the index and swap atomically
Rarely changesPeriodic full rebuild; gate the swap on a retrieval eval

Whatever the cadence, gate re-ingests on an eval so a chunker or model change cannot silently degrade retrieval, and store a build timestamp so “confident but wrong after a refresh” is diagnosable.

Evaluating retrieval separately from generation

A good answer can hide bad retrieval and vice versa, so measure them apart.

MetricMeasuresExample
Recall@kWas the gold passage in the top-k?8 of 10 queries had it in top-5 → 0.80
MRRRank of the first relevant hitRanks 1, 3, 2 → (1 + 1/3 + 1/2)/3 = 0.61
Precision@kFraction of top-k that are relevant2 of top-5 relevant → 0.40
FaithfulnessIs the answer supported by the retrieved context?47 of 50 claims grounded → 0.94
Answer relevanceDoes the answer address the question?Rubric or model-graded

Worked comparison on a 200-query set:

ConfigRecall@5MRRFaithfulnessAnswer relevance
Dense only0.710.520.860.83
Hybrid (RRF)0.830.610.900.86
Hybrid + reranker0.830.740.940.90

Reranking barely moves Recall@5 (same candidates) but lifts MRR and faithfulness sharply by putting the right passage first, where the model reads it earliest. Report every metric per segment (document type, source, language) — an aggregate 0.90 can hide one source at 0.40.

Failure taxonomy with fixes

SymptomLikely causeFix
Answer wrong, gold passage not retrievedChunk drift, wrong embedding model, poor queryVersion chunking; align embed models; add query rewriting
Gold passage retrieved but ranked lowNo reranker; bad fusion weightsAdd a cross-encoder reranker; tune RRF k
Exact codes / IDs missedDense-only retrievalAdd BM25 (hybrid)
Retrieved and top-ranked but answer wrongModel ignoring context; context not passedCheck the prompt actually includes chunks; constrain to context
Confident but wrong after a refreshStale or re-chunked indexRe-run retrieval evals; check build timestamp
One source consistently badAggregate hid itReport per-segment; fix that source’s ingest
Restricted content surfacedNo permission filter at retrievalFilter by identity at query time

Cost and latency budget worked example

A support assistant over 40,000 articles, target end-to-end latency under 2 seconds at p95, gpt-5.6-terra for generation.

text
Per query budget (p95 target: 1900 ms)
query rewrite (gpt-5.6-luna, low effort) ~120 ms
hybrid retrieve (dense + BM25, top_k 40) ~140 ms
cross-encoder rerank (40 -> 8) ~180 ms
generate grounded answer (terra, 8 chunks) ~1200 ms
citation assembly (code) ~20 ms
--------
total (p95) ~1660 ms ✓ under 1900

Cost levers, cheapest first: cut the number of chunks fed to the model (8 → 5 with a stronger reranker) before touching the model tier; use gpt-5.6-luna for the rewrite step; cache the stable system prefix on the API; and reserve a larger model for only the queries a router flags as hard. The wrong instinct — “use a bigger model to fix wrong answers” — spends money on a generation problem that is usually a retrieval problem.

Common misconceptions

MisconceptionRealityWhy it matters on the assessment
“More context always helps”Noise lowers faithfulness; rerank to a small top-nOver-retrieval distractor
“Dense retrieval covers everything”It misses exact IDs and codes; add BM25Hybrid-search signal
“Wrong answer means fix the prompt”Check retrieval first after a refreshRoot-cause distractor
“A prompt rule keeps restricted data out”Permissions are enforced at retrieval, not by promptPermission trap
“Fine-tune to add new facts”Fine-tuning is for style; RAG for factsRAG-vs-fine-tuning distractor
“Reranking improves recall”It improves ordering and precision, not recallMetric-confusion distractor

Scenario walkthrough

A team runs a KB assistant over 40,000 articles refreshed weekly. This week it began giving confident wrong answers. Users ask both natural-language questions and exact error codes. What is the FIRST step and the right architecture?

  1. FIRST step — pull traces and check whether the gold passage is even retrieved. It surfaced that this week’s re-ingest changed the chunker (drift). This is a retrieval problem, so rewriting the prompt would waste the cycle.
  2. Chunking — structural or parent-child for articles: child chunks for precise matches, parent sections for context; version the config and gate re-ingests on an eval.
  3. Retrieval — hybrid dense + BM25, because error codes need the exact-term matching that dense retrieval misses.
  4. Reranking — cross-encoder to top-8; the gold article was retrieved but sat at rank 14 before reranking.
  5. Permissions — filter by the asking user’s identity at query time, since some articles are internal-only.
  6. Eval — recall@k, MRR and faithfulness, reported per source, so a single failing source cannot hide behind the aggregate.

Rejected alternatives: rewriting the system prompt (retrieval was broken), fine-tuning on the KB (facts change weekly — that is RAG’s job), dense-only retrieval (misses error codes), and trusting the aggregate score (it masked the failing source).

Key takeaways

  • RAG for large or changing corpora with citations; long context for small stable corpora; fine-tuning for style, not facts.
  • Chunk to match the data shape; parent-child balances precise hits with context; version the config and gate re-ingests.
  • Keep the index and query on the same embedding model; re-embed the whole corpus on any model change.
  • Choose the file search tool for speed with less infrastructure, a self-managed store when you need chunking, hybrid retrieval, a reranker or permission integration you control.
  • Hybrid (dense + BM25) plus a cross-encoder reranker beats either alone; RRF fuses without score calibration.
  • Enforce permissions and freshness at retrieval time, and evaluate retrieval (recall@k, MRR) separately from generation (faithfulness), per segment.

Last updated Sep 18, 2026