पाठ 15 / 31

Documents, Loaders और Text Splitters

Files को Document objects में लोड करें और समझदारी से split करें।

Document = पाठ + metadata

LangChain Document में page_content (पाठ) और metadata (स्रोत, पृष्ठ, तिथि, अनुमतियाँ) होते हैं। Loaders files, web pages, databases और cloud drives को Documents में पढ़ते हैं। Text splitter लंबे दस्तावेज़ों को embedding और prompting के लिए उपयुक्त आकार के chunks में बाँटता है। RecursiveCharacterTextSplitter separators क्रम से आज़माता है (अनुच्छेद विराम, पंक्ति विराम, spaces) ताकि chunks प्राकृतिक सीमाओं पर समाप्त हों, chunk_size और chunk_overlap अक्षरों में (या tokens में, यदि आप token-आधारित लंबाई function दें)। आकार अपने सवालों पर परखें; बहुत छोटे chunks संदर्भ खोते हैं और बहुत बड़े विषय धुँधले करते हैं।

Split, embed, retrieve, उत्तर

दस्तावेज़ chunks बनते हैं, chunks vectors, और retrieved chunks prompt में जाते हैं।

चार चरण: split, embed, retrieve, उत्तर।
चित्र 4.1 — Split, embed, retrieve और उत्तर।

Overlap के साथ splitting, चलाकर

मैंने यह Python virtual environment में langchain-core 1.6.6, langchain-text-splitters 1.1.2 और llama-index-core 0.14.25 के साथ ऑफ़लाइन चलाया। किसी API key या network call की ज़रूरत नहीं क्योंकि असली मॉडल की जगह fake मॉडल या खिलौना embedding है। chunk_size=90 और chunk_overlap=20 के साथ पाठ 4 chunks बनता है। दूसरा chunk 86 अक्षरों पर कटता है और तीसरा ("keep context across boundaries.") दूसरे के अंत से दोहराए शब्दों से शुरू होता है: यही overlap है।

from langchain_text_splitters import RecursiveCharacterTextSplitter

text = ("LangChain splits long documents into chunks.\n\n"
        "Each chunk keeps some overlap with the previous one. "
        "This helps retrieval keep context across boundaries.\n\n"
        "Smaller chunks are more precise; larger chunks carry more context.")
splitter = RecursiveCharacterTextSplitter(chunk_size=90, chunk_overlap=20)
chunks = splitter.split_text(text)
for i, c in enumerate(chunks):
    print(i, len(c), repr(c))

Output:

0 44 'LangChain splits long documents into chunks.'
1 86 'Each chunk keeps some overlap with the previous one. This helps retrieval keep context'
2 31 'keep context across boundaries.'
3 66 'Smaller chunks are more precise; larger chunks carry more context.'

त्वरित जाँच: Chunk overlap किसलिए है?

  • Chunk सीमाओं के पार संदर्भ बनाए रखना
  • लागत दोगुनी करना
  • पाठ encrypt करना
  • Duplicates हटाना
Answer

Chunk सीमाओं के पार संदर्भ बनाए रखना — सीमाओं पर थोड़ा पाठ दोहराना किसी विचार को आधा कटने से बचाता है।