Lesson 2 / 27

The RAG Pipeline End to End

Name the offline and online stages and what each produces.

Offline indexing, online answering

A RAG system has two phases. Offline (indexing): load documents, clean them, split into chunks, attach metadata, compute embeddings and store chunks, vectors and metadata in an index. Online (query time): take the user's question, optionally rewrite it, retrieve the top chunks (keyword, vector or both), optionally rerank them, build a prompt with the chunks, call the model, then check and return the answer with citations. Each stage can fail separately, so you evaluate and debug them separately.

The pipeline at a glance

The two lines show the offline and online flows.

OFFLINE  documents -> clean -> chunk -> (+metadata) -> embed -> index
ONLINE   question -> [rewrite] -> retrieve top-k -> [rerank] -> prompt
                  -> LLM -> verify citations -> answer + sources

Log every stage

Save the rewritten query, retrieved chunk IDs, scores and the final prompt for each request. Debugging a bad answer is nearly impossible without them.

Quick check: Which step happens offline?

  • Chunking and embedding the documents
  • Answering the user
  • Reranking for this question
  • Verifying citations
Answer

Chunking and embedding the documents — Indexing prepares the knowledge base before any question arrives.