Lesson 1 / 27

Why RAG Exists

Understand the knowledge problems of plain language models.

Models do not know your data

A language model only knows what was in its training data, frozen at a knowledge cutoff. It has never seen your company handbook, yesterday's tickets or a private database, and when it lacks the facts it may produce a fluent but invented answer (hallucination). Retrieval-Augmented Generation (RAG) fixes this at question time: first retrieve the passages that are relevant to the question from your own documents, then augment the prompt with them, then let the model generate an answer based on that text. The facts live in your documents, where they can be updated and cited, not inside the model's weights.

Look up, then answer

RAG fetches relevant text from your documents and gives it to the model with the question.

Four steps: question, retrieve, augment, generate.
Figure 1.1 — Question, retrieve, augment and generate.

An open-book exam

A closed-book student relies on memory and may guess. An open-book student looks up the right page, then writes the answer and says which page it came from.

Start with the question, not the tool

Collect 20 real questions your users ask before choosing any vector database. They tell you what to retrieve.

Quick check: Where do the facts live in a RAG system?

  • In the tokenizer
  • Only in the model weights
  • In your documents, retrieved at question time
  • In the GPU
Answer

In your documents, retrieved at question time — RAG keeps knowledge outside the model so it can be updated and cited.