ML Atlas

08 · LLMs · 5 min read · Interactive · updated

What is RAG (retrieval-augmented generation) and how does it work?

In short

RAG retrieves document passages that match a question and pastes them into the prompt. The model answers from sources rather than from training memory alone.

What it is

RAG (retrieval-augmented generation) is an architecture in which a language model, before answering, receives document passages retrieved for the question. The system first searches a knowledge base (with a keyword search engine, a vector search, or both), picks a few of the most relevant passages and inserts them into the prompt with the instruction "answer based on the sources below". The name and the first version were proposed by Lewis et al. (2020).

Intuition: it is the difference between a closed-book exam and an open-book one. The model still has to understand the question and be able to formulate an answer, but it takes the facts from the pages it is pointed to, not from hazy memories of pretraining.

RAG is today the most common way to "connect" an LLM to company documentation, policies, a ticket database or current news — without fine-tuning the model.

Mechanism — why it works this way

Two memories. Knowledge stored in the model's weights (parametric memory) is broad but imprecise, frozen on the day training ended and impossible to verify. A document base (non-parametric memory) is exact, easy to update, and you can point to where a sentence came from. RAG combines the two: the model contributes language understanding, the base contributes facts.

Why it reduces hallucinations. The model generates what is probable in the given context. When the context contains a paragraph with the answer, its content becomes the most probable continuation — the model "copies and paraphrases" instead of guessing. This reduces fabrication but does not eliminate it: the model may combine passages incorrectly or ignore the source in favour of its own knowledge.

The pipeline. (1) Indexing: documents are split into chunks of a few hundred tokens, usually overlapping, and converted into embeddings or a word index. (2) Retrieval: the question goes through the same process, and the system returns the k nearest chunks. (3) Optional reranking: a more accurate but slower model scores question–chunk pairs and reorders the results. (4) Generation: the chunks are placed in the prompt, often with a request to cite the sources.

Retrieval is the bottleneck. If the right chunk does not make it into the context, even the best model will not answer correctly. Lexical search (BM25, TF-IDF) is good at catching proper names and codes, but understands neither synonyms nor inflection. Dense retrieval (embeddings, Karpukhin et al., 2020) captures meaning but can be weaker on rare identifiers. That is why hybrid search is popular.

Limitations. Poor chunking can cut an answer in half. Too many chunks distract the model, and information in the middle of a long context tends to be overlooked. RAG copes poorly with questions that require aggregating information across the whole collection ("how many contracts did we sign in 2023?") — that is a job for a database and SQL queries instead.

By example

The knowledge base holds four sentences from a Polish company handbook: (1) "Urlop wypoczynkowy przysługuje pracownikowi w wymiarze 20 lub 26 dni rocznie." ("An employee is entitled to 20 or 26 days of annual leave per year."), (2) "Wniosek o urlop składa się w systemie kadrowym co najmniej tydzień wcześniej." ("A leave request is submitted in the HR system at least a week in advance."), (3) one about parking, (4) one about business trips. The question: "Jak złożyć wniosek o urlop?" ("How do I submit a leave request?"). TF-IDF retrieval with cosine similarity gives: sentence 2 — 0.207, sentence 1 — 0.084, sentences 3 and 4 — 0. Chunk 2 comes out on top, and the model will answer from it that the request is submitted in the HR system a week in advance.

This small example also shows the weakness of lexical search in an inflected language such as Polish: "złożyć" (to submit) and "składa" (is submitted) are different words to TF-IDF, so the match rests only on "wniosek" (request) and "urlop" (leave). An embedding model would treat them as close. The scale of indexing is easy to compute: a 100,000-token manual cut into 500-token chunks with an overlap of 50 gives 223 chunks to index.

In practice

  • Chunks of 200–800 tokens with a 10–20% overlap are a reasonable starting point; split along the document's structure (headings, paragraphs), not in the middle of a sentence.
  • For embeddings, use a model that handles your language well (a multilingual one for non-English text); test it on your own questions.
  • Combine BM25 with vector search and add a reranker (a cross-encoder) for the top few dozen candidates.
  • Measure retrieval quality (is the right chunk in the top k, recall@k) and answer quality (faithfulness to the sources) separately.
  • Common mistake: judging the whole system only by its answers — when something goes wrong, you cannot tell whether the retriever or the model failed.

Frequently asked questions

RAG or fine-tuning?
RAG for facts that change, that must be cited, or that are confidential and scattered across documents. Fine-tuning for style, format and skills. The two are often combined: a fine-tuned model makes better use of the sources it is given.
Does RAG eliminate hallucinations?
No, but it clearly reduces them and makes them detectable, because the answer can be compared with the source. It is worth asking the model for quotations and for an "I don't know" answer when the sources do not contain the information.
Do I need a vector database?
With a few thousand chunks, a plain search over all vectors in memory is enough. A specialized database or an approximate index (e.g. HNSW) pays off with hundreds of thousands or millions of chunks.

Sources

  • Lewis P. et al., 2020, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", NeurIPS 2020.
  • Karpukhin V. et al., 2020, "Dense Passage Retrieval for Open-Domain Question Answering", EMNLP 2020.
  • Robertson S., Zaragoza H., 2009, "The Probabilistic Relevance Framework: BM25 and Beyond", Foundations and Trends in Information Retrieval 3(4).
  • Liu N. F. et al., 2024, "Lost in the Middle: How Language Models Use Long Contexts", Transactions of the ACL 12.
  • Manning C. D., Raghavan P., Schütze H., 2008, "Introduction to Information Retrieval", Cambridge University Press, ch. 6 (TF-IDF and the vector space model).

See also