Lesson 15 / 31

Documents, Loaders and Text Splitters

Load files into Document objects and split them sensibly.

Document = text + metadata

A LangChain Document holds page_content (text) and metadata (source, page, date, permissions). Loaders read files, web pages, databases and cloud drives into Documents. A text splitter divides long documents into chunks sized for embedding and prompting. RecursiveCharacterTextSplitter tries separators in order (paragraph breaks, line breaks, spaces) so chunks end at natural boundaries, with chunk_size and chunk_overlap in characters (or tokens, if you provide a token-based length function). Test sizes on your own questions; very small chunks lose context and very large ones blur topics.

Split, embed, retrieve, answer

Documents become chunks, chunks become vectors, and retrieved chunks go into the prompt.

Four stages: split, embed, retrieve, answer.
Figure 4.1 — Split, embed, retrieve and answer.

Splitting with overlap, run

I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. With chunk_size=90 and chunk_overlap=20, the text becomes 4 chunks. The second chunk is cut at 86 characters and the third ("keep context across boundaries.") starts with words repeated from the end of the second: that is the overlap.

from langchain_text_splitters import RecursiveCharacterTextSplitter

text = ("LangChain splits long documents into chunks.\n\n"
        "Each chunk keeps some overlap with the previous one. "
        "This helps retrieval keep context across boundaries.\n\n"
        "Smaller chunks are more precise; larger chunks carry more context.")
splitter = RecursiveCharacterTextSplitter(chunk_size=90, chunk_overlap=20)
chunks = splitter.split_text(text)
for i, c in enumerate(chunks):
    print(i, len(c), repr(c))

Output:

0 44 'LangChain splits long documents into chunks.'
1 86 'Each chunk keeps some overlap with the previous one. This helps retrieval keep context'
2 31 'keep context across boundaries.'
3 66 'Smaller chunks are more precise; larger chunks carry more context.'

Quick check: What is chunk overlap for?

  • Keeping context across chunk boundaries
  • Doubling the cost
  • Encrypting text
  • Deleting duplicates
Answer

Keeping context across chunk boundaries — Repeating a little text at boundaries avoids cutting an idea in half.