Lesson 5 / 27

Loading and Cleaning Documents

Extract clean text from PDFs, web pages and office files.

Garbage in, garbage out

Real documents are messy. PDFs have headers, footers, page numbers, two columns and tables that extract as scrambled text; web pages have menus and cookie banners; scanned files need OCR. Clean before indexing: remove boilerplate, fix encodings, keep headings and list structure, convert tables to a readable form (rows as text or Markdown), deduplicate repeated documents and drop obsolete versions. Always look at a sample of extracted text with your own eyes; many "retrieval" problems are really extraction problems.

Good input, good index

Clean text, sensible chunks and useful metadata decide retrieval quality.

Three steps: load, chunk, label.
Figure 2.1 — Load, chunk and label.

Keep the source with the text

Store the file name, page number, section title and URL with every piece of text from the very first step; you will need them for citations.

Quick check: Why look at samples of extracted text?

  • To reduce the font size
  • Extraction errors silently ruin retrieval
  • To train the model
  • To speed up the GPU
Answer

Extraction errors silently ruin retrieval — Scrambled tables or missing text cannot be retrieved correctly.