# The RAG Pipeline End to End — Retrieval-Augmented Generation (RAG)

Source: https://www.geekswithgeeks.com/en/rag/b-pipeline

> Name the offline and online stages and what each produces.

## Offline indexing, online answering

A RAG system has two phases. **Offline (indexing)**: load documents, clean them, split into **chunks**, attach **metadata**, compute **embeddings** and store chunks, vectors and metadata in an **index**. **Online (query time)**: take the user's question, optionally rewrite it, **retrieve** the top chunks (keyword, vector or both), optionally **rerank** them, build a **prompt** with the chunks, call the model, then check and return the answer with **citations**. Each stage can fail separately, so you evaluate and debug them separately.

## The pipeline at a glance

The two lines show the offline and online flows.

```text
OFFLINE  documents -> clean -> chunk -> (+metadata) -> embed -> index
ONLINE   question -> [rewrite] -> retrieve top-k -> [rerank] -> prompt
                  -> LLM -> verify citations -> answer + sources
```

## Log every stage

Save the rewritten query, retrieved chunk IDs, scores and the final prompt for each request. Debugging a bad answer is nearly impossible without them.

**Quiz:** Which step happens offline?

- [x] Chunking and embedding the documents
- [ ] Answering the user
- [ ] Reranking for this question
- [ ] Verifying citations

*Answer:* Chunking and embedding the documents. Indexing prepares the knowledge base before any question arrives.
