पाठ 5 / 29
मूल्यांकन Harness
सिस्टम का कोई भी संस्करण उन्हीं मामलों पर चलाएँ और श्रेणी के अनुसार परिणाम रिपोर्ट करें।
हर बार वही मामले
मूल्यांकन harness छोटा प्रोग्राम है जो test मामलों (इनपुट, साथ में अपेक्षित उत्तर या जाँच नियम) का set लेता है, आपके system का उम्मीदवार संस्करण हर पर चलाता है, आउटपुट score करता है, और विफल मामलों की सूची के साथ श्रेणी के अनुसार रिपोर्ट छापता है। इसे तेज़, जहाँ संभव निश्चित (random seeds तय, कम temperature), और एक command से कोई भी चला सके ऐसा बनाएँ, स्थानीय रूप से और CI में। मामले data files के रूप में रखें, उन्हें श्रेणियों (भाषाएँ, विषय, कठिनाई, विरोधी) से tag करें, किनारे के मामले और "मना करना चाहिए" मामले शामिल करें, और held-out set रखें जिसके विरुद्ध आप tune न करें। Scoring तरीक़े, सस्ते से महँगे: exact match या regex, programmatic जाँचें (वैध JSON, सही fields, सीमा में संख्याएँ), संदर्भ-आधारित समानता, LLM-as-judge, और मानव समीक्षा। जहाँ programmatic जाँचें आपकी परवाह पकड़ें वहाँ उन्हें प्राथमिकता दें।
मामले, scores, शोर, तुलना
Test set बनाएँ, स्वचालित रूप से score करें, सांख्यिकीय शोर समझें, और संस्करणों की निष्पक्ष तुलना करें।
दो prompt संस्करणों की तुलना, चलाकर
मैंने यह सादे Python 3 (सिर्फ़ standard library) से चलाया। यह गढ़ी संख्याओं और seeded random draws वाला अनुकरण है, इसलिए यह प्रभाव का रूप दिखाता है, किसी उत्पाद का माप नहीं। दो stand-in classifiers छह labelled मामलों पर "prompt v1" और "prompt v2" की भूमिका निभाते हैं। रिपोर्ट v1 के लिए 4/6 और v2 के लिए 6/6, प्रति श्रेणी pass गिनती और ठीक वे मामले दिखाती है जिनमें v1 विफल हुआ (ids 2 और 4)। असली harness मॉडल बुलाएगा; ढाँचा वही है।
import re
CASES = [
{"id": 1, "cat": "billing", "input": "I was charged twice", "check": ("equals", "billing")},
{"id": 2, "cat": "billing", "input": "refund my invoice", "check": ("equals", "billing")},
{"id": 3, "cat": "technical", "input": "app crashes on start", "check": ("equals", "technical")},
{"id": 4, "cat": "technical", "input": "cannot log in", "check": ("equals", "technical")},
{"id": 5, "cat": "other", "input": "where is your office", "check": ("equals", "other")},
{"id": 6, "cat": "other", "input": "tell me a joke", "check": ("equals", "other")},
]
def score(output, check):
kind, want = check
return output.strip().lower() == want if kind == "equals" else bool(re.search(want, output))
def prompt_v1(text): # stand-in for "model + prompt version 1": keyword rules only
t = text.lower()
return "billing" if "charged" in t else "technical" if "crash" in t else "other"
def prompt_v2(text): # version 2 adds a few more cues
t = text.lower()
if any(w in t for w in ("charged", "refund", "invoice")): return "billing"
if any(w in t for w in ("crash", "log in", "error")): return "technical"
return "other"
def run(model):
by_cat, results = {}, []
for c in CASES:
ok = score(model(c["input"]), c["check"]); results.append((c["id"], ok))
by_cat.setdefault(c["cat"], []).append(ok)
return results, {k: f"{sum(v)}/{len(v)}" for k, v in by_cat.items()}
for name, model in (("v1", prompt_v1), ("v2", prompt_v2)):
results, cats = run(model)
print(name, "pass", sum(ok for _, ok in results), "/", len(results), "| by category:", cats, "| failed ids:", [i for i, ok in results if not ok])
Output:
v1 pass 4 / 6 | by category: {'billing': '1/2', 'technical': '1/2', 'other': '2/2'} | failed ids: [2, 4]
v2 pass 6 / 6 | by category: {'billing': '2/2', 'technical': '2/2', 'other': '2/2'} | failed ids: []मामले data files में रखें
सादी JSONL files की समीक्षा, versioning और विस्तार टीम का कोई भी कर सकता है।
त्वरित जाँच: ऐसा held-out test set क्यों रखें जिसके विरुद्ध आप tune न करें?
- Held-out sets तेज़ चलते हैं
- उन्हीं मामलों पर tune करना गुणवत्ता को बढ़ाकर दिखाता है (test पर overfitting)
- APIs उन्हें माँगती हैं
- वे scoring की ज़रूरत हटाते हैं
Answer
उन्हीं मामलों पर tune करना गुणवत्ता को बढ़ाकर दिखाता है (test पर overfitting) — आपको ऐसे मामलों पर ईमानदार माप चाहिए जिन पर system tune नहीं हुआ।