पाठ 9 / 27

Loss और Perplexity

मापें कि language model पाठ कितना अच्छा अनुमानित करता है।

मॉडल कितना चकित है

प्रशिक्षण cross-entropy loss घटाता है: -log(मॉडल ने सही अगले token को जो संभावना दी) का औसत। कम बेहतर है। Perplexity exp(loss) है और इसे "मॉडल उतना ही अनिश्चित है मानो इतने tokens में समान रूप से चुन रहा हो" पढ़ा जा सकता है। Perplexity 1 मतलब सटीक अनुमान; 10 मतलब निष्पक्ष 10-तरफ़ा चुनाव जितना अनिश्चित। Perplexity एक ही पाठ और tokenizer पर मॉडलों की तुलना को उपयोगी है, पर कम मान अपने आप साबित नहीं करता कि मॉडल सहायक या सत्यवादी है।

Bigram model की perplexity, चलाकर

मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। प्रशिक्षण वाक्य 1.817 और दूसरा देखा वाक्य 2.048 पाता है। अनदेखी शब्द-जोड़ी वाला वाक्य अनंत पाता है, इसीलिए असली systems संभावनाएँ smooth करते हैं, और समान 10-तरफ़ा अनुमान ठीक 10 पाता है।

import math

text = "the cat sat on the mat . the dog sat on the log . the cat saw the dog ."
tokens = text.split()
from collections import defaultdict, Counter
counts = defaultdict(Counter)
for a, b in zip(tokens, tokens[1:]):
    counts[a][b] += 1

def prob(a, b):
    total = sum(counts[a].values())
    return counts[a][b] / total if total else 0.0

def perplexity(seq):
    logp = 0.0
    for a, b in zip(seq, seq[1:]):
        p = prob(a, b)
        logp += math.log(p) if p > 0 else float("-inf")
    return math.exp(-logp / (len(seq) - 1))

print("seen sentence  :", round(perplexity("the cat sat on the mat .".split()), 3))
print("other sentence:", round(perplexity("the cat saw the dog .".split()), 3))
print("unseen pair   :", perplexity("the sat cat on the mat .".split()))
print("uniform over 10:", round(math.exp(-math.log(1 / 10)), 3))

Output:

seen sentence  : 1.817
other sentence: 2.048
unseen pair   : inf
uniform over 10: 10.0

Perplexity की तुलना निष्पक्ष रखें

Perplexity की तुलना सिर्फ़ उसी पाठ और उसी tokenizer पर नापे मानों में करें; अलग tokenizers पाठ अलग बाँटते हैं और संख्याएँ बदल देते हैं।

त्वरित जाँच: Perplexity 1 का मतलब…

  • पाठ में एक token है
  • मॉडल कुछ नहीं जानता
  • मॉडल पाठ को सटीक अनुमानित करता है
  • मॉडल पक्षपाती है
Answer

मॉडल पाठ को सटीक अनुमानित करता है — Perplexity 1 हर सही अगले token को संभावना 1 देने के बराबर है।