Lesson 19 / 31

LlamaIndex Data Model: Documents and Nodes

Understand what loaders produce and how parsers split data.

Document in, Nodes out

In LlamaIndex a Document is a container for a source of data (a file, a page, a database row) with text and metadata. Data loaders (many are available through LlamaHub and integration packages) produce Documents. A node parser such as SentenceSplitter splits Documents into Node objects (chunks) that keep their metadata and relationships to the source Document and to neighbouring nodes. Nodes are the unit that gets embedded, stored and retrieved. Metadata (file name, page, date) travels with each node and can be used for filtering and citations.

From files to answers

Documents become nodes, nodes become an index, and a query engine answers from it.

Four objects: document, node, index, engine.
Figure 5.1 — Document, node, index and engine.

Splitting into nodes, run

I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. A 12-sentence document with chunk_size=60 becomes 3 nodes of 4 sentences each. Every node keeps the file metadata of its source Document.

from llama_index.core import Document
from llama_index.core.node_parser import SentenceSplitter

text = " ".join(f"Sentence {i} explains one idea about indexing data for retrieval." for i in range(1, 13))
doc = Document(text=text, metadata={"file": "notes.txt"})
nodes = SentenceSplitter(chunk_size=60, chunk_overlap=0).get_nodes_from_documents([doc])
for n in nodes:
    sentences = n.get_content().count("Sentence")
    print(n.metadata["file"], "| sentences in node:", sentences, "| first words:", " ".join(n.get_content().split()[:3]))
print("nodes:", len(nodes))

Output:

notes.txt | sentences in node: 4 | first words: Sentence 1 explains
notes.txt | sentences in node: 4 | first words: Sentence 5 explains
notes.txt | sentences in node: 4 | first words: Sentence 9 explains
nodes: 3

Quick check: What does a Node represent in LlamaIndex?

  • A prompt template
  • A server in a cluster
  • A GPU core
  • A chunk of a Document, with metadata and relationships
Answer

A chunk of a Document, with metadata and relationships — Nodes are the retrievable units produced from Documents.