पाठ 9 / 27
Loss और Perplexity
मापें कि language model पाठ कितना अच्छा अनुमानित करता है।
मॉडल कितना चकित है
प्रशिक्षण cross-entropy loss घटाता है: -log(मॉडल ने सही अगले token को जो संभावना दी) का औसत। कम बेहतर है। Perplexity exp(loss) है और इसे "मॉडल उतना ही अनिश्चित है मानो इतने tokens में समान रूप से चुन रहा हो" पढ़ा जा सकता है। Perplexity 1 मतलब सटीक अनुमान; 10 मतलब निष्पक्ष 10-तरफ़ा चुनाव जितना अनिश्चित। Perplexity एक ही पाठ और tokenizer पर मॉडलों की तुलना को उपयोगी है, पर कम मान अपने आप साबित नहीं करता कि मॉडल सहायक या सत्यवादी है।
Bigram model की perplexity, चलाकर
मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। प्रशिक्षण वाक्य 1.817 और दूसरा देखा वाक्य 2.048 पाता है। अनदेखी शब्द-जोड़ी वाला वाक्य अनंत पाता है, इसीलिए असली systems संभावनाएँ smooth करते हैं, और समान 10-तरफ़ा अनुमान ठीक 10 पाता है।
import math
text = "the cat sat on the mat . the dog sat on the log . the cat saw the dog ."
tokens = text.split()
from collections import defaultdict, Counter
counts = defaultdict(Counter)
for a, b in zip(tokens, tokens[1:]):
counts[a][b] += 1
def prob(a, b):
total = sum(counts[a].values())
return counts[a][b] / total if total else 0.0
def perplexity(seq):
logp = 0.0
for a, b in zip(seq, seq[1:]):
p = prob(a, b)
logp += math.log(p) if p > 0 else float("-inf")
return math.exp(-logp / (len(seq) - 1))
print("seen sentence :", round(perplexity("the cat sat on the mat .".split()), 3))
print("other sentence:", round(perplexity("the cat saw the dog .".split()), 3))
print("unseen pair :", perplexity("the sat cat on the mat .".split()))
print("uniform over 10:", round(math.exp(-math.log(1 / 10)), 3))
Output:
seen sentence : 1.817 other sentence: 2.048 unseen pair : inf uniform over 10: 10.0
Perplexity की तुलना निष्पक्ष रखें
Perplexity की तुलना सिर्फ़ उसी पाठ और उसी tokenizer पर नापे मानों में करें; अलग tokenizers पाठ अलग बाँटते हैं और संख्याएँ बदल देते हैं।
त्वरित जाँच: Perplexity 1 का मतलब…
- पाठ में एक token है
- मॉडल कुछ नहीं जानता
- मॉडल पाठ को सटीक अनुमानित करता है
- मॉडल पक्षपाती है
Answer
मॉडल पाठ को सटीक अनुमानित करता है — Perplexity 1 हर सही अगले token को संभावना 1 देने के बराबर है।