Lesson 2 / 27
The RAG Pipeline End to End
Name the offline and online stages and what each produces.
Offline indexing, online answering
A RAG system has two phases. Offline (indexing): load documents, clean them, split into chunks, attach metadata, compute embeddings and store chunks, vectors and metadata in an index. Online (query time): take the user's question, optionally rewrite it, retrieve the top chunks (keyword, vector or both), optionally rerank them, build a prompt with the chunks, call the model, then check and return the answer with citations. Each stage can fail separately, so you evaluate and debug them separately.
The pipeline at a glance
The two lines show the offline and online flows.
OFFLINE documents -> clean -> chunk -> (+metadata) -> embed -> index
ONLINE question -> [rewrite] -> retrieve top-k -> [rerank] -> prompt
-> LLM -> verify citations -> answer + sourcesLog every stage
Save the rewritten query, retrieved chunk IDs, scores and the final prompt for each request. Debugging a bad answer is nearly impossible without them.
Quick check: Which step happens offline?
- Chunking and embedding the documents
- Answering the user
- Reranking for this question
- Verifying citations
Answer
Chunking and embedding the documents — Indexing prepares the knowledge base before any question arrives.