# Loading and Cleaning Documents — Retrieval-Augmented Generation (RAG)

Source: https://www.geekswithgeeks.com/en/rag/i-load

> Extract clean text from PDFs, web pages and office files.

## Garbage in, garbage out

Real documents are messy. PDFs have headers, footers, page numbers, two columns and tables that extract as scrambled text; web pages have menus and cookie banners; scanned files need **OCR**. Clean before indexing: remove boilerplate, fix encodings, keep **headings and list structure**, convert tables to a readable form (rows as text or Markdown), **deduplicate** repeated documents and drop obsolete versions. Always look at a sample of extracted text with your own eyes; many "retrieval" problems are really extraction problems.

## Good input, good index

Clean text, sensible chunks and useful metadata decide retrieval quality.

![Three steps: load, chunk, label.](assets/figures/rag/section-2-map.svg) — Figure 2.1 — Load, chunk and label.

## Keep the source with the text

Store the file name, page number, section title and URL with every piece of text from the very first step; you will need them for citations.

**Quiz:** Why look at samples of extracted text?

- [ ] To reduce the font size
- [x] Extraction errors silently ruin retrieval
- [ ] To train the model
- [ ] To speed up the GPU

*Answer:* Extraction errors silently ruin retrieval. Scrambled tables or missing text cannot be retrieved correctly.
