Lesson 19 / 31
LlamaIndex Data Model: Documents and Nodes
Understand what loaders produce and how parsers split data.
Document in, Nodes out
In LlamaIndex a Document is a container for a source of data (a file, a page, a database row) with text and metadata. Data loaders (many are available through LlamaHub and integration packages) produce Documents. A node parser such as SentenceSplitter splits Documents into Node objects (chunks) that keep their metadata and relationships to the source Document and to neighbouring nodes. Nodes are the unit that gets embedded, stored and retrieved. Metadata (file name, page, date) travels with each node and can be used for filtering and citations.
From files to answers
Documents become nodes, nodes become an index, and a query engine answers from it.
Splitting into nodes, run
I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. A 12-sentence document with chunk_size=60 becomes 3 nodes of 4 sentences each. Every node keeps the file metadata of its source Document.
from llama_index.core import Document
from llama_index.core.node_parser import SentenceSplitter
text = " ".join(f"Sentence {i} explains one idea about indexing data for retrieval." for i in range(1, 13))
doc = Document(text=text, metadata={"file": "notes.txt"})
nodes = SentenceSplitter(chunk_size=60, chunk_overlap=0).get_nodes_from_documents([doc])
for n in nodes:
sentences = n.get_content().count("Sentence")
print(n.metadata["file"], "| sentences in node:", sentences, "| first words:", " ".join(n.get_content().split()[:3]))
print("nodes:", len(nodes))
Output:
notes.txt | sentences in node: 4 | first words: Sentence 1 explains notes.txt | sentences in node: 4 | first words: Sentence 5 explains notes.txt | sentences in node: 4 | first words: Sentence 9 explains nodes: 3
Quick check: What does a Node represent in LlamaIndex?
- A prompt template
- A server in a cluster
- A GPU core
- A chunk of a Document, with metadata and relationships
Answer
A chunk of a Document, with metadata and relationships — Nodes are the retrievable units produced from Documents.