पाठ 6 / 27

Chunking रणनीतियाँ

पाठ को सही आकार और overlap के साथ retrieve-योग्य टुकड़ों में बाँटें।

न बहुत बड़ा, न बहुत छोटा

Retrieval chunks पर चलता है, इसलिए उनका आकार मायने रखता है। बहुत छोटा हो तो chunk संदर्भ खो देता है ("यह", "ऊपर वाला") और पूरा उत्तर नहीं रख पाता। बहुत बड़ा हो तो उसका embedding कई विषयों को धुँधला कर देता है, retrieval कम सटीक होता है और prompts महँगे। आम शुरुआती बिंदु 200 से 500 tokens और 10 से 20% overlap हैं, पर सर्वोत्तम आकार आपके दस्तावेज़ों और सवालों पर निर्भर है, इसलिए परखें। संरचना-जागरूक chunking आम तौर पर निश्चित windows से बेहतर है: headings, अनुच्छेदों या वाक्यों पर बाँटें, तालिका या code block को पूरा रखें, और हर chunk के आगे दस्तावेज़ व अनुभाग का शीर्षक जोड़ें ताकि वह अकेले अर्थपूर्ण रहे।

निश्चित windows बनाम वाक्य chunks, चलाकर

मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। 15 overlap वाली निश्चित 60-अक्षर windows शब्दों को आधा काटती हैं ("retriev", "nks are"), जबकि वाक्य-आधारित chunking पूरे वाक्य साथ रखती है।

text = ("RAG splits documents into chunks. Each chunk is embedded. "
        "At question time the nearest chunks are retrieved. "
        "They are placed in the prompt. The model answers from them.")

def fixed(text, size, overlap):
    step = size - overlap
    return [text[i:i + size] for i in range(0, len(text), step) if text[i:i + size].strip()]

def by_sentence(text, max_chars):
    out, cur = [], ""
    for s in text.split(". "):
        s = s.strip().rstrip(".") + "."
        if cur and len(cur) + len(s) + 1 > max_chars:
            out.append(cur); cur = s
        else:
            cur = (cur + " " + s).strip()
    if cur: out.append(cur)
    return out

for c in fixed(text, 60, 15): print(repr(c))
print("---")
for c in by_sentence(text, 70): print(repr(c))

Output:

'RAG splits documents into chunks. Each chunk is embedded. At'
'is embedded. At question time the nearest chunks are retriev'
'nks are retrieved. They are placed in the prompt. The model '
'mpt. The model answers from them.'
---
'RAG splits documents into chunks. Each chunk is embedded.'
'At question time the nearest chunks are retrieved.'
'They are placed in the prompt. The model answers from them.'

असली सवालों पर chunk आकार परखें

2 या 3 आकार आज़माएँ और तयशुदा सवालों पर recall की तुलना करें (evaluation खंड देखें)। अनुमान न लगाएँ।

त्वरित जाँच: बहुत बड़े chunks की एक समस्या क्या है?

  • वे metadata हटाते हैं
  • वे रखे नहीं जा सकते
  • वे हमेशा सटीकता बढ़ाते हैं
  • उनके embeddings कई विषय धुँधला करते हैं और prompts महँगे हो जाते हैं
Answer

उनके embeddings कई विषय धुँधला करते हैं और prompts महँगे हो जाते हैं — बड़े chunks retrieval की सटीकता घटाते हैं और prompt tokens बर्बाद करते हैं।