012 min
The problem RAG solves
A language model keeps what it learned during training inside its own weights. The 2020 paper that named RAG calls this parametric memory. The facts are spread across billions of numbers, and no single number can be pointed at (Lewis et al., NeurIPS 2020).
Three limits follow from that.
The facts stop at the training cut-off. A support assistant trained before a price change will state the old price, with the same confidence as a correct one.
Nothing can be checked. The model produced a sentence. There is no passage behind that sentence for a reader to open.
Changing one fact means training the model again. That is slow and expensive, and it does not work well. Ovadia et al. at Microsoft (December 2023) compared the two methods on knowledge the model had never seen. Retrieval beat unsupervised fine-tuning at getting new facts into a model's answers.
The approach before RAG that did produce citations was extractive question answering. A retriever found candidate passages, and a second model marked the exact span of text that answered the question. Karpukhin et al. built one of these in 2020. It could point at a source, but copying a span was all it could do. It could not join two passages into one answer, restate a technical paragraph in plain words, or reply that the documents do not cover the question.
RAG was proposed to get both properties in one system: an answer written fresh for the question, and a passage the reader can open.
022 min
How RAG works
Follow one question the whole way: a customer asks a product's support assistant, "Does the Pro plan include audit logs?" The documents are that product's own help centre.
Before any question arrives
The help centre is split into chunks β short passages, each a few hundred words. Anthropic's published setup from September 2024 used chunks of about 800 tokens, a token being roughly three quarters of an English word.
Each chunk is turned into an embedding: a long list of numbers that represents the chunk's meaning. Two passages about the same subject get similar lists. The embeddings go into a vector store, which is storage built to find the closest lists of numbers to a given list, fast, across millions of entries.
When the question arrives
The question goes through the same embedding model. The vector store returns the top-k chunks β the k closest matches, where k is a number you choose. Anthropic tested 5, 10 and 20 and reported 20 as the best of the three for their setup.
Most systems run a second search beside it. BM25 scores by exact word overlap, the way a classic search box does. It catches audit logs as a phrase, where an embedding search may instead return a chunk about activity history. Running both and merging the two result lists is called hybrid retrieval.
The merged shortlist often passes through reranking. A slower, more accurate model reads the question and each candidate chunk together, scores them again, and puts the best three or five first.
Then the prompt is assembled β the surviving chunks, the question, and an instruction that limits the model to those chunks.
Answer only from the passages below. If they do not contain the answer, reply that the documentation does not cover it.
The model writes the answer and names the chunk it used, so the customer can open that help page.
The step that decides the rest
The model never reads the help centre. It reads whatever retrieval handed it. If the audit-logs paragraph ranked eleventh and k was ten, the model is being asked a question whose answer is not in its prompt. It will usually answer anyway.
So every quality problem in a RAG system is either a retrieval problem or a generation problem, and the two need different fixes. That split is the most useful fact to remember about the architecture.
Lewis et al. call the index non-parametric memory β knowledge kept outside the weights. Replace December's index with January's and the system's knowledge changes the same afternoon, with no training run.
032 min
Where RAG came from
The term comes from one paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. It was written by Patrick Lewis and eleven co-authors at Facebook AI Research, University College London and New York University. It was posted to arXiv on 22 May 2020 and published at NeurIPS 2020.
Two pieces already existed. REALM (Guu et al., February 2020) had shown that a language model could be pre-trained with a retriever attached to it. Dense Passage Retrieval, or DPR (Karpukhin et al., April 2020, EMNLP 2020), had shown something else. A retriever trained on question-and-passage pairs beat BM25 keyword search by 9 to 19 percentage points on top-20 passage accuracy. That measure is the share of questions where the answer appears somewhere in the first twenty passages returned.
Lewis et al. connected a DPR retriever to BART, a pre-trained model that writes text, and trained the two together. The index was a December 2018 copy of Wikipedia, cut into 21 million passages of 100 words each. They tested two versions. RAG-Sequence picks one set of passages and uses it for the whole answer. RAG-Token can use a different passage for each word it writes.
The name was not laboured over. Patrick Lewis told NVIDIA in January 2025 that the team had always planned a nicer sounding name. Nobody had a better idea by the time the paper was written. Lewis added: "We definitely would have put more thought into the name had we known our work would become so widespread."
042 min
RAG's first benchmark results
The 2020 paper's main table is open-domain question answering. The model gets a question and no passage, and is scored by exact match: the answer string it writes has to match a human-written answer exactly. Natural Questions (NQ) is real Google search queries answered from Wikipedia. TriviaQA (TQA), WebQuestions (WQ) and CuratedTrec (CT) are three other question sets. All four numbers below are test-set exact match, from Table 1 of Lewis et al. (2020).
| Model | NQ | TQA | WQ | CT |
|---|---|---|---|---|
| T5-11B (no retrieval) | 34.5 | β | 37.4 | β |
| T5-11B+SSM (no retrieval) | 36.6 | β | 44.7 | β |
| REALM | 40.4 | β | 40.7 | 46.8 |
| DPR | 41.5 | 57.9 | 41.1 | 50.6 |
| RAG-Token | 44.1 | 55.2 | 45.5 | 50.0 |
| RAG-Sequence | 44.5 | 56.8 | 45.2 | 52.2 |
Read the first row against the last. T5-11B has 11 billion parameters and answers only from its weights; it scores 34.5 on Natural Questions. RAG-Sequence scores 44.5 using BART-large, a generator with 400 million parameters, plus the Wikipedia index. Ten points of exact match, from a far smaller model, because the facts were fetched rather than memorised.
Two cautions before quoting these. They are 2020 numbers against a December 2018 index, and models released since score much higher on all four sets. They show that retrieval helped in 2020, not what any system scores today. Exact match is also a harsh measure: an answer that is correct but worded differently scores zero.
052 min
RAG inside a shipped legal product
The benchmark above tested a research system on an encyclopedia, where answers are short and the corpus is clean. The second case changes both variables: a paid product, a professional corpus, and answers long enough that the citation matters as much as the text.
Legal research vendors sold RAG as the fix for invented case law. Magesh et al., at Stanford's RegLab and Institute for Human-Centered AI, ran the first preregistered evaluation of those products and published it in May 2024. They had to test from the outside, because the systems are closed.
The vendors' claims, quoted in the paper, were that RAG "eliminat[es]" hallucinations (Casetext, 2023), "avoid[s]" them (Thomson Reuters, 2023), and produces "hallucination-free" legal citations (LexisNexis, 2023).
The measured rates: Lexis+ AI and Ask Practical Law AI each produced incorrect information more than 17% of the time. Westlaw's AI-Assisted Research was wrong more than 34% of the time. The same study found these tools did make fewer errors than general-purpose GPT-4, which the authors call a substantial improvement.
Comparing the two cases gives the lesson. On a clean corpus with short answers, adding retrieval raised the score a lot. On a professional corpus with long answers, adding retrieval lowered the error rate but still left roughly one answer in six, or one in three, wrong.
Retrieval limits what the model can say. It does not force the model to stay inside that limit. Barnett et al. (2024) name this failure separately from retrieval failure. The answer is present in the retrieved text, and the model still fails to extract it or answers at the wrong level of specificity.
062 min
RAG vs. fine-tuning and long context
Four neighbours get mixed up with RAG, and one fact separates each of them from it.
| Approach | What it changes | Choose it when | Main limit |
|---|---|---|---|
| AIRAG | |||
| AIFine-tuning | |||
| Not in the library yetLong-context prompting | |||
| Not in the library yetSemantic search | |||
| AITool use |
The deciding fact between RAG and fine-tuning: fine-tuning changes behaviour, retrieval changes available facts. Ovadia et al. (2023) tested this and found retrieval better for new knowledge. Teams often need both, and the order matters β fine-tune for the output format you want, retrieve for the facts that go inside it.
The deciding fact between RAG and long context is the size of the corpus. Anthropic's guidance from September 2024 is direct: under about 200,000 tokens, roughly 500 pages, put the whole knowledge base in the prompt and skip retrieval. Above that, the documents no longer fit and retrieval is needed.
071 min
What RAG costs to run
Four costs, and only the first is a one-off.
Indexing. Every document is embedded once. Change the chunk size or swap the embedding model and the whole corpus is re-embedded. Anthropic reported in September 2024 that generating one extra sentence of context per chunk cost $1.02 per million document tokens, using prompt caching. That is the price of one optional improvement, not of indexing itself.
Tokens per question. This is the cost people miss. Retrieval makes every prompt bigger. Twenty chunks of 800 tokens is about 16,000 extra input tokens on every question asked, whether or not the answer needed them. That was the configuration Anthropic reported as best in their September 2024 tests.
Latency. Each query runs an embedding call, a vector search and then generation. Reranking adds another model call before generation. Anthropic's own note on it is that reranking "inevitably adds a small amount of latency", and that the trade is accuracy for speed.
Upkeep. The index becomes out of date as documents change, so re-indexing is a scheduled job, not a launch task. The evaluation set needs the same care. Barnett et al. (2024) concluded from three deployed systems that a RAG system can only really be validated while it is running.
082 min
What RAG does not solve
Barnett et al. (2024) built RAG systems for three domains β research, education and biomedical β and published the seven ways they broke:
- Missing content β the corpus does not contain the answer at all
- Missed top-ranked documents β the answer exists but ranks below the cut-off
- Not in context β retrieved, then dropped when the passages were combined
- Not extracted β present in the prompt, and the model still misses it
- Wrong format β a table or list was asked for and ignored
- Incorrect specificity β answered too broadly or too narrowly to be useful
- Incomplete β correct as far as it goes, missing available detail
Only the first three are retrieval problems. The rest happen after the right text is already in the prompt, which is why adding a better retriever does not fix them.
Three situations call for a different tool entirely.
Questions about the whole corpus. "What are the main themes across these documents?" cannot be answered by fetching the twenty best-matching passages, because no passage contains the answer. Edge et al. at Microsoft (April 2024) built GraphRAG for exactly this gap, summarising a corpus into a graph first rather than searching it.
Questions with an arithmetic answer. Counting, totalling or filtering belongs in a database query, not in retrieved text.
Untrusted documents. Anything retrieved is read by the model as input, and Greshake et al. (2023) showed that attackers can plant instructions in content a system is likely to fetch. A retrieved page can carry text aimed at the model rather than at the reader. This attack is called indirect prompt injection, and retrieval is what makes it reachable.
092 min
What RAG changes in your work
Engineers. Measure retrieval and generation separately, from the first week. Two questions, asked of every wrong answer: was the correct passage in the prompt, and did the model use it? The fixes do not overlap. For the first, change the chunk size, add hybrid retrieval, or change k. For the second, change the instruction, change the model, or send fewer chunks. Position matters too. Liu et al. (TACL, 2023) found that models use information best at the start or the end of a long input, and worst in the middle. So the order chunks are placed in is a real setting, not a detail.
Product managers. Your evaluation set needs questions the documents cannot answer. Without them you never measure whether the system refuses, and refusing correctly is most of what separates a usable assistant from a confident one. Second habit: when an answer is wrong, ask where the correct text ranked before blaming the model.
Designers. A citation has to open the passage itself, not a 40-page document. Opening the source is the only way a reader can check an answer. Treat "not in the documentation" as a designed state with its own layout and a next step, not as an error message.
Founders. A RAG feature cannot be better than the documents you already own. If the help centre is thin, contradictory or out of date, retrieval will find exactly that and the model will repeat it. Budget for writing and maintaining that corpus alongside the engineering.
102 min
Where the RAG evidence is contested
Two findings disagree with normal RAG practice, and neither is settled.
Cuconasu et al. (2024) examined what kind of passages should go in the prompt. Two results went against the usual advice. Passages the retriever scored highly but that do not contain the answer lowered accuracy. And adding genuinely random documents to the prompt improved accuracy, by up to 35% in their tests. The authors present this as a reason to study retrieval strategy rather than as a recommendation, and it has not been widely replicated, so treat it as contested. The practical reading is narrower: tuning only for top-k relevance may be the wrong target, because near-miss passages are the expensive mistake.
Li et al. at Google (EMNLP 2024, industry track) compared RAG against long-context models directly. Their finding: given enough resources, long context beat RAG on average across the public datasets they tested, on the three most recent models available to them. RAG's advantage was cost. Their hybrid method, Self-Route, sends each query to whichever path suits it and cut cost by 65% for Gemini-1.5-Pro and 39% for GPT-4o against long context alone.
Both results are tied to particular datasets and to 2024 model versions, and in this field both go out of date quickly. Two things survive for a practitioner. Run the comparison on your own corpus rather than inheriting a conclusion. Keep retrieval where you need citations or predictable cost, even where long context scores higher.
111 min
How RAG changed since 2020
The original RAG trained things. Lewis et al. fine-tuned the generator and the retriever's question encoder together. Their own ablations on the dev set show how much that mattered. Replacing the trained dense retriever with BM25 dropped Natural Questions exact match from 43.5 to 29.7. Freezing the retriever instead of training it dropped the same score to 37.8.
Almost nothing called RAG today trains anything. The normal build is an off-the-shelf embedding model, a vector store, a frozen commercial language model, and a prompt assembled at query time. That is a different system using the original's name, which is worth knowing when reading the 2020 paper's results.
The word then split into named variants. Gao et al. published a survey in December 2023, revised in March 2024, that divided the field into Naive RAG, Advanced RAG and Modular RAG. Microsoft published GraphRAG in April 2024 for whole-corpus questions. Anthropic published contextual retrieval in September 2024, which adds a sentence of surrounding context to each chunk before the chunk is embedded. Anthropic reported that this cut top-20 retrieval failures by 49% when paired with contextual BM25, and by 67% with reranking added.
So "RAG" now names a family of systems, not the method in the paper.
?8 questions
Questions people ask
Is RAG the same as fine-tuning?
Does RAG need a vector database?
Does RAG stop hallucinations?
When should I use long context instead of RAG?
How much does RAG cost to run?
Does RAG require GPUs or training data?
Why does RAG miss answers that are in my documents?
Can retrieved documents attack a RAG system?
Β§14 sources
Sources on RAG
Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020.
Karpukhin, V. et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. EMNLP 2020.
Guu, K. et al. (2020). REALM: Retrieval-Augmented Language Model Pre-Training.
Ovadia, O. et al. (2023). Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs.
Show all 14 sourcesShow fewer sources
Liu, N. F. et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. TACL, volume 11.
Barnett, S. et al. (2024). Seven Failure Points When Engineering a Retrieval Augmented Generation System.
Magesh, V. et al. (2024). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Stanford RegLab and HAI. summary at
Cuconasu, F. et al. (2024). The Power of Noise: Redefining Retrieval for RAG Systems.
Li, Z. et al. (2024). Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach. EMNLP 2024 industry track.
Edge, D. et al. (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. Microsoft.
Greshake, K. et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.
Gao, Y. et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey.
Anthropic (19 September 2024). Introducing Contextual Retrieval.
Merritt, R. (31 January 2025). What Is Retrieval-Augmented Generation, aka RAG? NVIDIA β source of the Patrick Lewis quotes.





