# Documents, Loaders and Text Splitters — LangChain / LlamaIndex

Source: https://www.geekswithgeeks.com/en/langchain-llamaindex/r-docs

> Load files into Document objects and split them sensibly.

## Document = text + metadata

A LangChain **`Document`** holds `page_content` (text) and `metadata` (source, page, date, permissions). **Loaders** read files, web pages, databases and cloud drives into Documents. A **text splitter** divides long documents into chunks sized for embedding and prompting. **`RecursiveCharacterTextSplitter`** tries separators in order (paragraph breaks, line breaks, spaces) so chunks end at natural boundaries, with `chunk_size` and `chunk_overlap` in characters (or tokens, if you provide a token-based length function). Test sizes on your own questions; very small chunks lose context and very large ones blur topics.

## Split, embed, retrieve, answer

Documents become chunks, chunks become vectors, and retrieved chunks go into the prompt.

![Four stages: split, embed, retrieve, answer.](assets/figures/langchain-llamaindex/section-4-map.svg) — Figure 4.1 — Split, embed, retrieve and answer.

## Splitting with overlap, run

I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. With `chunk_size=90` and `chunk_overlap=20`, the text becomes 4 chunks. The second chunk is cut at 86 characters and the third ("keep context across boundaries.") starts with words repeated from the end of the second: that is the overlap.

```python
from langchain_text_splitters import RecursiveCharacterTextSplitter

text = ("LangChain splits long documents into chunks.\n\n"
        "Each chunk keeps some overlap with the previous one. "
        "This helps retrieval keep context across boundaries.\n\n"
        "Smaller chunks are more precise; larger chunks carry more context.")
splitter = RecursiveCharacterTextSplitter(chunk_size=90, chunk_overlap=20)
chunks = splitter.split_text(text)
for i, c in enumerate(chunks):
    print(i, len(c), repr(c))

```

Output:

```
0 44 'LangChain splits long documents into chunks.'
1 86 'Each chunk keeps some overlap with the previous one. This helps retrieval keep context'
2 31 'keep context across boundaries.'
3 66 'Smaller chunks are more precise; larger chunks carry more context.'
```

**Quiz:** What is chunk overlap for?

- [x] Keeping context across chunk boundaries
- [ ] Doubling the cost
- [ ] Encrypting text
- [ ] Deleting duplicates

*Answer:* Keeping context across chunk boundaries. Repeating a little text at boundaries avoids cutting an idea in half.
