पाठ 9 / 29

Context Window बजट के रूप में

जो समाए उसमें सबसे प्रासंगिक सामग्री चुनें।

सब कुछ एक ही जगह के लिए होड़ करता है

Context window पाठ की वह अधिकतम मात्रा (tokens में) है जिसे मॉडल एक call में देख सकता है। उसमें निर्देश, tool विवरण, अब तक की बातचीत (हर tool परिणाम समेत), agent द्वारा पढ़ी files और उत्तर के लिए जगह समानी चाहिए। लंबे contexts धीमे और महँगे होते हैं, और मॉडल बहुत लंबे context में दबी जानकारी कम भरोसे से उपयोग करते हैं, इसलिए ज़्यादा हमेशा बेहतर नहीं। इसलिए harness context चुनता है: वह files को कार्य से प्रासंगिकता (नाम, symbols, खोज hits, हाल के संपादन) से क्रमित करता है, token बजट के भीतर शीर्ष वाली शामिल करता है, और पूरी files की जगह केवल प्रासंगिक पंक्ति-सीमाएँ दिखाता है। आप मदद कर सकते हैं: ज़रूरी files के नाम बताकर, समान मौजूदा कोड की ओर इशारा करके, और कार्यों को एक बदलाव पर केंद्रित रखकर।

छोटी window, सोच-समझकर चुनी

Agent सिर्फ़ वही जानता है जो उसके context में है, इसलिए वहाँ क्या रखा गया और वह कितना सुव्यवस्थित है, गुणवत्ता तय करता है।

चार औज़ार: pack, निर्देश, compact, सौंपना।
चित्र 3.1 — Pack, निर्देश, compact और सौंपना।

Token बजट में प्रासंगिक files भरना, चलाकर

मैंने यह सादे Python 3 (सिर्फ़ standard library) से चलाया, अस्थायी folder में बनाए throwaway project के साथ। Files को इससे score किया जाता है कि query शब्द कितनी बार आते हैं (file नाम तीन गुना गिने जाते हैं)। Pricing file, उसके tests और cart 300 में से 150 tokens में समा जाते हैं; लंबा असंबंधित history दस्तावेज़ शून्य score पाता है और छोड़ा जाता है। असली harnesses समृद्ध संकेत उपयोग करते हैं, पर विचार वही है।

def estimate_tokens(text): return max(1, len(text) // 4)          # rough: 4 characters per token

def pack_context(files, query_terms, budget):
    """Pick the most relevant files that fit the token budget."""
    def score(item):
        name, text = item
        return sum(text.lower().count(t) + 3 * name.lower().count(t) for t in query_terms)
    chosen, used = [], 0
    for name, text in sorted(files.items(), key=lambda kv: -score(kv)):
        cost = estimate_tokens(text)
        if score((name, text)) == 0 or used + cost > budget:
            continue
        chosen.append((name, cost)); used += cost
    return chosen, used

files = {
    "shop/pricing.py": "def apply_discount(price, percent): ... discount discount tax " * 5,
    "shop/cart.py": "class Cart: total discount " * 4,
    "docs/history.md": "release notes and history " * 200,
    "tests/test_pricing.py": "def test_discount ... discount " * 6,
}
chosen, used = pack_context(files, ["discount", "pricing"], budget=300)
print("chosen:", chosen)
print("tokens used:", used, "of 300")

Output:

chosen: [('shop/pricing.py', 77), ('tests/test_pricing.py', 46), ('shop/cart.py', 27)]
tokens used: 150 of 300

ज़रूरी files के नाम बताएँ

Agent को बताना कि कौन-सी files या functions शामिल हैं, खोज चरण और context बचाता है।

त्वरित जाँच: पूरा repository हमेशा context में क्यों न रखें?

  • क्योंकि context मुफ़्त है
  • Repositories पढ़े नहीं जा सकते
  • वह शायद ही समाता है, ज़्यादा ख़र्चता है, धीमा है और प्रासंगिक हिस्से दबा देता है
  • क्योंकि मॉडल कोड नापसंद करते हैं
Answer

वह शायद ही समाता है, ज़्यादा ख़र्चता है, धीमा है और प्रासंगिक हिस्से दबा देता है — प्रासंगिक सामग्री चुनना सब कुछ भेजने से बेहतर है।