पाठ 21 / 29
Benchmarks और उनकी सीमाएँ
समझें कि issue-resolution benchmarks क्या मापते हैं और क्या चूकते हैं।
उपयोगी, पर आपका codebase नहीं
SWE-bench जैसे सार्वजनिक benchmarks agent को असली repository और असली GitHub issue देते हैं और जाँचते हैं कि परिणामी patch से project के छिपे tests पास होते हैं या नहीं। यह systems की तुलना का मूल्यवान तरीक़ा है, और scores तेज़ी से बढ़े हैं। पर सावधानी रखें: कार्य खुले-स्रोत Python projects के छोटे समूह से आते हैं, tests अधूरे हो सकते हैं (patch पास होकर भी ग़लत या अधूरा हो सकता है), कार्य प्रशिक्षण डेटा में आ चुके हो सकते हैं, scores सिर्फ़ मॉडल पर नहीं बल्कि harness, prompts और प्रयासों की संख्या पर बहुत निर्भर हैं, और वे आपकी भाषा, परंपराओं, build system या कोड-गुणवत्ता मानकों के बारे में कम बताते हैं। अपने निर्णयों के लिए छोटा आंतरिक benchmark बनाएँ: अपने repository के 20 से 50 असली पुराने कार्य (ठीक किए bugs, छोटी सुविधाएँ) ज्ञात शुरुआती commit और सफलता तय करने वाले tests के साथ, फिर उस पर tools और settings की तुलना करें।
ईमानदारी से मापें
Benchmarks, pass@k, run की लागत और असली उत्पादकता संख्याएँ हर एक कहानी का हिस्सा बताती हैं।
आंतरिक benchmark डिज़ाइन करना
नुस्ख़ा; संख्याएँ अपनी टीम के अनुसार ढालें।
1. collect 20-50 closed issues/PRs that had a clear fix and tests (bugs, small features)
2. for each: record the commit BEFORE the fix, the task text a human would give, and the tests that define success
3. run each tool/setting on every task in a clean checkout, same time/step/cost limits
4. score: tests pass? (automatic) + diff reviewed for scope, quality, safety (human, sampled)
5. report: success rate, median cost and time per task, failure categories
6. re-run when the model, tool version, prompt file or settings changeनिजी holdout रखें
अपने आंतरिक benchmark कार्य प्रकाशित न करें; कुछ अनदेखे रखें ताकि tuning overfit न करे।
त्वरित जाँच: सार्वजनिक benchmarks के अलावा आंतरिक benchmark क्यों बनाएँ?
- सार्वजनिक benchmarks का उपयोग अवैध है
- सार्वजनिक benchmarks आपकी भाषा, परंपराओं और कोड-गुणवत्ता मानक को न दर्शाएँ
- आंतरिक benchmarks को tests नहीं चाहिए
- Scores कभी मायने नहीं रखते
Answer
सार्वजनिक benchmarks आपकी भाषा, परंपराओं और कोड-गुणवत्ता मानक को न दर्शाएँ — आपके workflow के निर्णयों को अपने कार्यों से प्रमाण चाहिए।