# Test Set और Harness बनाना — Prompt Engineering

Source: https://www.geekswithgeeks.com/hi/prompt-engineering/v-testset

> तयशुदा उदाहरणों पर score के साथ prompt संस्करणों की तुलना करें।

## धारणाओं की जगह संख्याएँ

अपेक्षित आउटपुट या ग्रेडिंग नियमों के साथ 30 से 200 **वास्तविक परीक्षण मामले** जुटाएँ, जिनमें सामान्य मामले, किनारे के मामले, विरोधी इनपुट और ऐसे उदाहरण हों जहाँ सही उत्तर "unknown" है। छोटा **harness** लिखें जो किसी prompt संस्करण को सब मामलों पर चलाकर score और विफलताएँ रिपोर्ट करे। जहाँ संभव हो **exact या नियम-आधारित जाँचें** उपयोग करें (label अपेक्षित के बराबर, JSON वैध, आवश्यक fields मौजूद) और खुले-अंत पाठ के लिए **rubrics या judge मॉडल**, मनुष्यों से नमूना-जाँच के साथ। एक बार में एक चीज़ बदलें और पूरा set दोबारा चलाएँ, क्योंकि एक मामले का सुधार अक्सर दूसरे को तोड़ देता है।

## भेजने से पहले मापें

Test set, versioning और लागत ट्रैकिंग prompting को engineering बनाते हैं।

![चार आदतें: परीक्षण, versioning, लागत, निगरानी।](assets/figures/prompt-engineering/section-7-map.svg) — चित्र 7.1 — परीक्षण, versioning, लागत और निगरानी।

## दो prompts की तुलना, चलाकर

मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। दो सरल functions दो prompt संस्करणों की जगह हैं। संस्करण A 4 में से 3 सही पाता है और "Refund my invoice" पर विफल होता है; संस्करण B, जो "refund" और "invoice" भी जानता है, 4 में से 4 पाता है। Harness ठीक दिखाता है कि कौन-सा मामला बदला।

```python
cases = [
    ("I was charged twice", "billing"),
    ("App crashes on start", "technical"),
    ("Where is your office", "other"),
    ("Refund my invoice", "billing"),
]

def prompt_a(text):          # keyword-only "model" standing in for prompt A
    t = text.lower()
    return "billing" if "charged" in t else "technical" if "crash" in t else "other"

def prompt_b(text):          # an improved version that also knows "refund"/"invoice"
    t = text.lower()
    if any(w in t for w in ("charged", "refund", "invoice")): return "billing"
    return "technical" if "crash" in t else "other"

for name, fn in (("A", prompt_a), ("B", prompt_b)):
    ok = [fn(t) == y for t, y in cases]
    print(name, f"{sum(ok)}/{len(cases)}", "failed:", [c[0] for c, o in zip(cases, ok) if not o])

```

Output:

```
A 3/4 failed: ['Refund my invoice']
B 4/4 failed: []
```

## हर production विफलता जोड़ें

असली user को बुरा उत्तर मिले तो वह इनपुट test set में जोड़ें ताकि वह चुपचाप कभी न बिगड़े।

**Quiz:** हर prompt बदलाव के बाद पूरा test set क्यों दोबारा चलाएँ?

- [ ] Versioning से बचने के लिए
- [ ] Prompts लंबे करने के लिए
- [x] एक मामले का सुधार दूसरों को तोड़ सकता है
- [ ] Tokenizer के लिए यह ज़रूरी है

*Answer:* एक मामले का सुधार दूसरों को तोड़ सकता है. Regression testing अनपेक्षित दुष्प्रभाव पकड़ता है।
