पाठ 24 / 29
Test Set और Harness बनाना
तयशुदा उदाहरणों पर score के साथ prompt संस्करणों की तुलना करें।
धारणाओं की जगह संख्याएँ
अपेक्षित आउटपुट या ग्रेडिंग नियमों के साथ 30 से 200 वास्तविक परीक्षण मामले जुटाएँ, जिनमें सामान्य मामले, किनारे के मामले, विरोधी इनपुट और ऐसे उदाहरण हों जहाँ सही उत्तर "unknown" है। छोटा harness लिखें जो किसी prompt संस्करण को सब मामलों पर चलाकर score और विफलताएँ रिपोर्ट करे। जहाँ संभव हो exact या नियम-आधारित जाँचें उपयोग करें (label अपेक्षित के बराबर, JSON वैध, आवश्यक fields मौजूद) और खुले-अंत पाठ के लिए rubrics या judge मॉडल, मनुष्यों से नमूना-जाँच के साथ। एक बार में एक चीज़ बदलें और पूरा set दोबारा चलाएँ, क्योंकि एक मामले का सुधार अक्सर दूसरे को तोड़ देता है।
भेजने से पहले मापें
Test set, versioning और लागत ट्रैकिंग prompting को engineering बनाते हैं।
दो prompts की तुलना, चलाकर
मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। दो सरल functions दो prompt संस्करणों की जगह हैं। संस्करण A 4 में से 3 सही पाता है और "Refund my invoice" पर विफल होता है; संस्करण B, जो "refund" और "invoice" भी जानता है, 4 में से 4 पाता है। Harness ठीक दिखाता है कि कौन-सा मामला बदला।
cases = [
("I was charged twice", "billing"),
("App crashes on start", "technical"),
("Where is your office", "other"),
("Refund my invoice", "billing"),
]
def prompt_a(text): # keyword-only "model" standing in for prompt A
t = text.lower()
return "billing" if "charged" in t else "technical" if "crash" in t else "other"
def prompt_b(text): # an improved version that also knows "refund"/"invoice"
t = text.lower()
if any(w in t for w in ("charged", "refund", "invoice")): return "billing"
return "technical" if "crash" in t else "other"
for name, fn in (("A", prompt_a), ("B", prompt_b)):
ok = [fn(t) == y for t, y in cases]
print(name, f"{sum(ok)}/{len(cases)}", "failed:", [c[0] for c, o in zip(cases, ok) if not o])
Output:
A 3/4 failed: ['Refund my invoice'] B 4/4 failed: []
हर production विफलता जोड़ें
असली user को बुरा उत्तर मिले तो वह इनपुट test set में जोड़ें ताकि वह चुपचाप कभी न बिगड़े।
त्वरित जाँच: हर prompt बदलाव के बाद पूरा test set क्यों दोबारा चलाएँ?
- Versioning से बचने के लिए
- Prompts लंबे करने के लिए
- एक मामले का सुधार दूसरों को तोड़ सकता है
- Tokenizer के लिए यह ज़रूरी है
Answer
एक मामले का सुधार दूसरों को तोड़ सकता है — Regression testing अनपेक्षित दुष्प्रभाव पकड़ता है।