API Developer Path
D5 · Retrieval-Augmented Generation
Chunking, embeddings, vector stores, file search, hybrid retrieval, citation formatting, grounded-answer evaluation, and handling freshness and permissions in retrieval on the OpenAI API.
This domain is about 16% of the OAI-API mock – roughly 10 of 60 items – and mirrors the Academy Build with Retrieval-Augmented Generation course (100 min). It tests whether you can ground a model’s answers in your own documents: chunk and embed content, store it in a vector store, retrieve with file search and hybrid methods, format citations, evaluate whether answers are actually grounded, and handle freshness and permissions so retrieval returns the right, allowed, current evidence.
What you need to know
RAG puts your documents in front of the model at answer time instead of relying on training knowledge. You chunk documents into passages, embed them into vectors, and store them in a vector store. At query time you retrieve the most relevant chunks (semantic, keyword, or hybrid) and pass them as context so the model answers from evidence and cites its sources. RAG is the right choice when knowledge is large, private, or changing – you update the store instead of retraining. The hard parts are not the plumbing: they are chunking well, retrieving the right passages, proving answers are grounded, keeping content fresh, and enforcing permissions so users only see what they may.
Learning objectives
By the end of this page you should be able to:
- Chunk documents sensibly and explain why chunk size matters.
- Create embeddings and a vector store, and answer with file search.
- Use hybrid retrieval (semantic + keyword) and know when it beats pure semantic.
- Format citations and evaluate grounded answers.
- Handle freshness so stale content is not retrieved.
- Enforce permissions so retrieval respects who may see what.
5.1 When RAG, and the pipeline
INGEST (offline) QUERY (online)──────────────── ──────────────documents user question │ chunk │ embed query ▼ ▼chunks ── embed ──► vectors ──► VECTOR STORE ──► retrieve top-k chunks │ ▼ prompt = question + chunks │ ▼ model answers WITH citationsRAG beats fine-tuning and beats stuffing everything in the prompt when the knowledge base is large, private, or frequently updated: you edit the store, not the model.
Assessment signal
“Answer only from our documents”, “cite the source”, “the knowledge changes weekly”, or “the answer must be grounded” are RAG items. “Fine-tune on the documents” is the classic distractor when the content changes or must be cited.
5.2 Chunking
Chunk size trades recall against precision and cost.
| Chunk size | Effect | Risk |
|---|---|---|
| Too large | Fewer, broad chunks; more tokens per hit | Retrieves irrelevant text; dilutes the signal |
| Too small | Precise but fragmented | Splits an answer across chunks; loses context |
| Right-sized (often a few hundred tokens, with overlap) | Coherent passages that each stand alone | — |
Overlap between adjacent chunks preserves context that would otherwise be cut at a boundary. Respect document structure – chunk on headings/paragraphs rather than fixed character counts where possible.
5.3 Embeddings and the vector store
Embeddings turn text into vectors so semantically similar passages sit near each other. A vector store indexes them for fast nearest-neighbour retrieval. With the OpenAI platform you can create a managed vector store and use file search without running your own vector database.
vs = client.vector_stores.create(name="hr-policies")client.vector_stores.files.upload_and_poll( vector_store_id=vs.id, file=open("leave-policy.pdf", "rb"),)
resp = client.responses.create( model="gpt-5.6-terra", input="How many days of paid parental leave do I get?", tools=[{"type": "file_search", "vector_store_ids": [vs.id]}],)print(resp.output_text) # answer grounded in the uploaded files, with citations5.4 Hybrid retrieval
Pure semantic search can miss exact terms (a product code, an error string, a legal clause number); pure keyword search misses paraphrases. Hybrid retrieval combines both and usually reranks.
| Method | Strong at | Weak at |
|---|---|---|
| Semantic (vector) | Paraphrase, meaning, synonyms | Exact IDs, rare tokens, exact phrases |
| Keyword (lexical) | Exact terms, codes, names | Synonyms, intent |
| Hybrid + rerank | Both, with a reranker to order | More moving parts to tune |
If users search by exact identifiers (SKUs, error codes, clause numbers), hybrid retrieval materially beats pure semantic.
5.5 Citations and grounded answers
A grounded answer points to the passage that supports each claim. Two things to get right:
- Citation formatting: return the source (file, section, and ideally the quoted span) so a reader can verify. File search returns citations you surface to the user.
- Grounded-answer discipline: instruct the model to answer only from retrieved context and to say “not found in the provided material” when the context does not support an answer – rather than filling the gap from training knowledge.
Answer: "Paid parental leave is 16 weeks." [leave-policy.pdf, §4.2]Check: open §4.2 → confirms 16 weeks → grounded.If the chunk said nothing → the model should say "not in the policy",not invent a number.5.6 Evaluating grounded answers
RAG has two failure surfaces, and you evaluate each:
| Failure | Where | Metric |
|---|---|---|
| Retrieval miss | Right chunk not retrieved | Retrieval recall @ k |
| Ungrounded answer | Chunk retrieved but answer not supported | Faithfulness / grounding score (LLM judge or human) |
| Incomplete answer | Some supporting chunks retrieved, others missed | Completeness against reference |
This is where D5 meets D3: build a dataset of questions with reference answers and supporting passages, and grade both retrieval and grounding. “The answer sounds right” is not a grounding measurement.
5.7 Freshness and permissions
| Concern | Problem | Handling |
|---|---|---|
| Freshness | Retrieving a superseded policy version | Re-index on change; store version/effective dates; filter to current; delete stale docs |
| Permissions | User retrieves a document they may not see | Enforce access at retrieval: filter by the user’s entitlements; partition stores by tenant/role |
Permissions belong in retrieval, not the prompt
Telling the model “don’t reveal documents the user can’t see” is not a control – the safe design filters the candidate set to what the user is entitled to before retrieval, so restricted content never reaches the context.
Decision framework
Use the GROUND framework to design and defend a RAG system.
| Letter | Step | Question |
|---|---|---|
| G | Good chunks | Are chunks coherent, right-sized, structure-aware, overlapped? |
| R | Retrieval method | Semantic, keyword, or hybrid – do users search by exact IDs? |
| O | Only from context | Does the model answer only from retrieved evidence and say “not found” otherwise? |
| U | Uphold permissions | Is the candidate set filtered to the user’s entitlements before retrieval? |
| N | New content | Is the store re-indexed on change with version/freshness filtering? |
| D | Demonstrate grounding | Do you evaluate retrieval recall and answer faithfulness on a dataset? |
The step most teams skip is D – they ship RAG that retrieves plausibly but never measure whether answers are actually supported.
Common mistakes
| Mistake | Why it happens | What to do instead |
|---|---|---|
| Fine-tuning on documents that change or must be cited | Confusing knowledge with retrieval | Use RAG; update the store, cite the source |
| Chunks too large | Fewer files to manage | Right-size with overlap; respect structure |
| Pure semantic search for exact-ID queries | Semantic is the default | Use hybrid retrieval when exact terms matter |
| Letting the model answer from training when context is empty | Helpfulness bias | Instruct “answer only from context; say not found” |
| No citations | Skipped for speed | Return source + span so answers are verifiable |
| Never measuring grounding | “It sounds right” | Evaluate retrieval recall and faithfulness on a dataset |
| Stale content retrieved | No re-index on change | Re-index on update; filter to current versions |
| Permissions enforced in the prompt | Easiest to bolt on | Filter candidates by entitlement before retrieval |
Scenario challenge
Scenario. A healthcare provider builds an assistant that answers clinician questions from internal protocols. Requirements: answers must cite the exact protocol section; clinicians must only retrieve protocols for their own department; protocols are revised monthly and the old version must never be quoted; and some questions reference exact protocol codes like PROT-CARD-014. An engineer proposes: fine-tune a model on all protocols, tell it in the system prompt to “only answer for the user’s department and cite sections”, and re-fine-tune monthly.
Expert reasoning trace.
- Fine-tuning is the wrong core. Protocols change monthly and answers must cite exact sections. Fine-tuning bakes a snapshot into weights, cannot cite a source span, and forces a costly retrain each month. RAG is correct: update the store, retrieve, cite.
- Permissions must be enforced at retrieval. A prompt instruction to “only answer for the user’s department” is not a control – the model still sees other departments’ chunks if they are in the candidate set. The safe design filters the candidate set to the clinician’s department before retrieval, e.g. partitioned stores or metadata filters, so out-of-scope protocols never reach context.
- Freshness needs versioning, not just re-indexing. Re-index monthly, but also tag each chunk with an effective date/version and filter retrieval to the current version so a superseded protocol is never quoted.
- Exact codes demand hybrid retrieval.
PROT-CARD-014is a rare exact token that pure semantic search may miss; hybrid (keyword + semantic) with reranking retrieves it reliably. - Ground and cite. Instruct the model to answer only from retrieved chunks and to return the protocol and section span; if the chunks do not answer, say so rather than inventing.
- Measure it. Build an eval of clinician questions with the correct protocol section, and grade retrieval recall and answer faithfulness – especially that no cross-department leakage occurs.
Exam-correct decision: RAG with a vector store, retrieval-time department filtering, version/freshness filtering, hybrid retrieval for exact codes, grounded-and-cited answers, and a grounding eval. Not fine-tuning, not prompt-based permission “controls”.
Assessment traps
| Trap | Why it is tempting | The discriminator |
|---|---|---|
| “Fine-tune on the documents” | Sounds like teaching the model | Changing/citable knowledge belongs in RAG, not weights |
| “Tell the model not to reveal restricted docs” | Simple to write | Permissions must filter the candidate set before retrieval |
| “Semantic search handles everything” | It is the default | Exact IDs/codes need hybrid retrieval |
| “It sounds right, so it’s grounded” | Fluency feels like grounding | Measure faithfulness against the retrieved passage |
| “Bigger chunks are safer” | Fewer to manage | Oversized chunks dilute relevance; right-size with overlap |
| “Re-indexing handles freshness” | Half-true | You also need version/date filtering so old versions are not quoted |
| “Put the whole knowledge base in the prompt” | Big context windows exist | Retrieval is cheaper, current, and permission-filterable |
Practice questions
Each item states how many responses to select. Commit before revealing.
Q1 · A knowledge base of 50,000 internal documents changes weekly and answers must cite sources. What is the BEST approach? (Select one)
A. Fine-tune a model on all documents weekly. B. RAG: chunk and embed into a vector store, retrieve at query time, and cite the retrieved sources. C. Paste all documents into every prompt. D. Rely on the model’s training knowledge.
Answer: B. Large, changing, citable knowledge is the canonical RAG case: update the store and cite retrieved passages. Weekly fine-tuning (A) is costly and cannot cite, pasting everything (C) is infeasible, and training knowledge (D) is neither private nor current.
Q2 · Users frequently search by exact error codes like `ERR-4021`. Pure semantic retrieval keeps missing them. What is the fix? (Select one)
A. Increase chunk size. B. Use hybrid retrieval that combines keyword and semantic search, with reranking. C. Fine-tune on the error codes. D. Lower the reasoning effort.
Answer: B. Exact rare tokens are a keyword strength and a semantic weakness, so hybrid retrieval catches them. Bigger chunks (A) do not fix lexical matching, fine-tuning (C) is the wrong tool, and effort (D) is unrelated to retrieval.
Q3 · Clinicians must only retrieve protocols for their own department. Where must this be enforced? (Select one)
A. In the system prompt, instructing the model not to reveal other departments. B. At retrieval, by filtering the candidate set to the user’s entitlements before passing context to the model. C. In the model’s training data. D. Nowhere; users self-police.
Answer: B. Permissions must filter what is retrieved so restricted content never reaches the context. A prompt instruction (A) still exposes chunks that are retrieved, training data (C) cannot enforce per-user access, and self-policing (D) is not a control.
Q4 · A retrieved chunk does not contain the answer, but the model responds with a confident number from its training knowledge. How do you prevent this? (Select one)
A. Increase the top-k retrieved chunks to 100. B. Instruct the model to answer only from retrieved context and to say the answer is not in the provided material otherwise. C. Use a larger model. D. Remove citations.
Answer: B. Grounded-answer discipline requires answering only from context and admitting when it is absent. Flooding with chunks (A) adds noise, a larger model (C) can still ungrounded-answer, and removing citations (D) makes it worse.
Q5 · Which TWO metrics should a grounded-answer eval measure? (Select two)
A. Retrieval recall – was the correct supporting chunk retrieved? B. Faithfulness – is each claim supported by the retrieved context? C. The model’s parameter count. D. The colour of the citation badge. E. Time of day of the query.
Answer: A and B. RAG has two failure surfaces – retrieval and grounding – so you measure retrieval recall and answer faithfulness. Parameter count (C), UI colour (D) and query time (E) are irrelevant to grounding quality.
Q6 · Chunks are set to 4,000 tokens each and answers now include lots of irrelevant text. What is the MOST likely cause? (Select one)
A. Chunks are too small. B. Chunks are too large, so each retrieved chunk dilutes the relevant passage with unrelated text. C. The embedding model is wrong. D. The vector store is full.
Answer: B. Oversized chunks retrieve broad passages that mix relevant and irrelevant text. They are not too small (A); nothing indicates the embedding model (C) or a full store (D) as the cause.
Q7 · Protocols are revised monthly and the previous version must never be quoted. Re-indexing alone is not enough. What else is needed? (Select one)
A. A larger context window. B. Version/effective-date metadata on chunks with retrieval filtered to the current version, and removal of superseded docs. C. Higher reasoning effort. D. More subagents.
Answer: B. Freshness needs version/date tagging and filtering so only the current version is retrievable, plus removing old ones. Context size (A), effort (C) and subagents (D) do not stop stale versions from being quoted.
Q8 · Why is overlap added between adjacent chunks? (Select one)
A. To increase storage cost deliberately. B. To preserve context that would otherwise be cut at a chunk boundary so an answer spanning the boundary is retrievable. C. To make embeddings faster. D. Overlap is never used.
Answer: B. Overlap keeps boundary-spanning context intact so a relevant passage is not split unrecoverably. It is not about cost (A) or speed (C), and it is a standard technique (D).
Q9 · A team ships RAG and checks quality by reading a few answers that 'sound right'. What is missing? (Select one)
A. Nothing; sounding right is sufficient. B. A grounded-answer evaluation measuring retrieval recall and faithfulness on a labelled dataset. C. A bigger vector store. D. Lower temperature.
Answer: B. Sounding right is not grounding; you must measure retrieval and faithfulness on a dataset. Reading a few answers (A) misses systematic failures, store size (C) is unrelated, and temperature (D) does not measure grounding.
Q10 · What does file search on a managed vector store give you out of the box? (Select one)
A. Nothing; you must build your own vector database. B. Retrieval over uploaded files with citations, without running your own vector database. C. Automatic fine-tuning of the model. D. Guaranteed EU data residency.
Answer: B. File search retrieves over a managed vector store and returns citations without you operating a vector DB. It does not require your own DB (A), does not fine-tune (C), and does not by itself guarantee residency (D).
Q11 · Which TWO are strong reasons to prefer RAG over fine-tuning for a knowledge task? (Select two)
A. The knowledge changes frequently and you want to update it without retraining. B. Answers must cite the exact source passage. C. You want to change the model’s writing style permanently. D. You want to reduce the number of API calls to zero. E. You want the model to memorise everything.
Answer: A and B. RAG suits changing knowledge (update the store) and citable answers (retrieve the passage). Permanent style change (C) is a fine-tuning use, RAG does not remove API calls (D), and memorisation (E) is the opposite of retrieval.
Q12 · A grounded assistant sometimes cites the wrong section for a correct-sounding answer. What is the appropriate diagnostic step? (Select one)
A. Assume the citation feature is broken and remove citations.
B. Run failure analysis on cited vs. supporting passages: check whether retrieval returned the right chunk and whether the model attributed the claim to the correct one.
C. Increase reasoning effort to max.
D. Switch to Chat Completions.
Answer: B. Wrong citations are a grounding/attribution failure diagnosed by comparing retrieved chunks to the claim; failure analysis isolates whether retrieval or attribution is at fault. Removing citations (A) hides the problem, effort (C) does not fix attribution, and the API surface (D) is irrelevant.
Q13 · A developer wants to 'just put the whole 900-page manual in the prompt every request' since the context window is 1.05M tokens. Why is RAG usually better here? (Select one)
A. RAG is always more accurate regardless of context. B. Retrieving only relevant chunks is cheaper per request, keeps answers focused, and lets you filter by permissions and freshness. C. The context window is actually too small for the manual. D. Prompts cannot contain documents.
Answer: B. Retrieval sends only relevant passages, cutting token cost and noise while enabling permission and freshness filtering the full-manual approach cannot. RAG is not unconditionally more accurate (A), 900 pages likely fit (C), and prompts can contain documents (D) – it is just wasteful.
Q14 · Retrieval recall is high but faithfulness is low in your eval. What does this indicate? (Select one)
A. The right chunks are retrieved but the model is not answering from them faithfully; tighten the grounding instruction and check attribution. B. The vector store is empty. C. Chunks are too small. D. The embedding model is broken.
Answer: A. High recall with low faithfulness means retrieval works but the model drifts from the evidence, so the fix is grounding discipline and attribution checks. An empty store (B) would tank recall too, and neither small chunks (C) nor a broken embedder (D) matches high recall.
Key takeaways
- RAG grounds answers in your documents; prefer it over fine-tuning when knowledge is large, private, changing or must be cited.
- Chunk into coherent, right-sized, overlapped passages that respect document structure.
- Store embeddings in a vector store and answer with file search; use hybrid retrieval when exact IDs matter.
- Make the model answer only from retrieved context and cite the source span; have it say “not found” rather than invent.
- Evaluate both retrieval recall and answer faithfulness on a labelled dataset.
- Enforce permissions by filtering the candidate set before retrieval, never with a prompt instruction.
- Keep content fresh with version/date filtering and re-indexing so superseded documents are never quoted.
Last updated Sep 18, 2026