पाठ 19 / 31
LlamaIndex Data Model: Documents और Nodes
समझें कि loaders क्या बनाते हैं और parsers डेटा कैसे split करते हैं।
Document अंदर, Nodes बाहर
LlamaIndex में Document डेटा के स्रोत (file, पृष्ठ, database पंक्ति) का पाठ और metadata वाला कंटेनर है। Data loaders (कई LlamaHub और integration packages से उपलब्ध) Documents बनाते हैं। SentenceSplitter जैसा node parser Documents को Node ऑब्जेक्ट्स (chunks) में बाँटता है जो अपना metadata और स्रोत Document व पड़ोसी nodes से संबंध रखते हैं। Nodes वह इकाई हैं जो embed, store और retrieve होती है। Metadata (file नाम, पृष्ठ, तिथि) हर node के साथ चलता है और filtering व citations में काम आता है।
Files से उत्तरों तक
Documents nodes बनते हैं, nodes index, और query engine उससे उत्तर देता है।
Nodes में splitting, चलाकर
मैंने यह Python virtual environment में langchain-core 1.6.6, langchain-text-splitters 1.1.2 और llama-index-core 0.14.25 के साथ ऑफ़लाइन चलाया। किसी API key या network call की ज़रूरत नहीं क्योंकि असली मॉडल की जगह fake मॉडल या खिलौना embedding है। chunk_size=60 के साथ 12-वाक्य का दस्तावेज़ 4-4 वाक्यों के 3 nodes बनता है। हर node अपने स्रोत Document का file metadata रखता है।
from llama_index.core import Document
from llama_index.core.node_parser import SentenceSplitter
text = " ".join(f"Sentence {i} explains one idea about indexing data for retrieval." for i in range(1, 13))
doc = Document(text=text, metadata={"file": "notes.txt"})
nodes = SentenceSplitter(chunk_size=60, chunk_overlap=0).get_nodes_from_documents([doc])
for n in nodes:
sentences = n.get_content().count("Sentence")
print(n.metadata["file"], "| sentences in node:", sentences, "| first words:", " ".join(n.get_content().split()[:3]))
print("nodes:", len(nodes))
Output:
notes.txt | sentences in node: 4 | first words: Sentence 1 explains notes.txt | sentences in node: 4 | first words: Sentence 5 explains notes.txt | sentences in node: 4 | first words: Sentence 9 explains nodes: 3
त्वरित जाँच: LlamaIndex में Node क्या दर्शाता है?
- Prompt template
- Cluster में server
- GPU core
- Document का chunk, metadata और संबंधों सहित
Answer
Document का chunk, metadata और संबंधों सहित — Nodes Documents से बनी retrieve होने वाली इकाइयाँ हैं।