पाठ 8 / 27
चलाने योग्य छोटा Language Model
अगला-शब्द अनुमान शुरू से अंत तक देखने को bigram मॉडल बनाएँ।
गिनती भी language model है
Bigram model प्रशिक्षण पाठ में गिनकर कि कौन-सा शब्द किसके बाद कितनी बार आता है, सिर्फ़ पिछले शब्द से अगला अनुमानित करता है। यह असली, भले छोटा, language model है: गिनतियाँ संभावनाएँ बनती हैं और पाठ बनाना उनसे sampling का लूप है। LLMs गिनती की तालिका को ऐसे neural network से बदलते हैं जो हज़ारों पिछले tokens देखता है, पर interface वही है: पिछला पाठ अंदर, अगले token पर वितरण बाहर।
गिनें और बनाएँ, चलाकर
मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। "the" के बाद मॉडल ने cat दो बार, dog दो बार, mat एक और log एक बार देखा। Seed 3 के साथ उसने "the cat saw the dog . the cat saw" बनाया।
import random
from collections import defaultdict, Counter
text = "the cat sat on the mat . the dog sat on the log . the cat saw the dog ."
tokens = text.split()
counts = defaultdict(Counter)
for a, b in zip(tokens, tokens[1:]):
counts[a][b] += 1
print("after 'the':", dict(counts["the"]))
print("after 'sat':", dict(counts["sat"]))
random.seed(3)
word, out = "the", ["the"]
for _ in range(8):
nxt = counts[word]
word = random.choices(list(nxt), list(nxt.values()))[0]
out.append(word)
print(" ".join(out))
Output:
after 'the': {'cat': 2, 'mat': 1, 'dog': 2, 'log': 1}
after 'sat': {'on': 2}
the cat saw the dog . the cat sawn-grams की दीवार क्यों
Bigram एक पिछला शब्द देखता है, इसलिए पहले का सब भूल जाता है। लंबे n-grams को घातीय रूप से ज़्यादा डेटा चाहिए। Neural networks इसके बजाय समान संदर्भों से सामान्यीकरण करते हैं।
त्वरित जाँच: Bigram model अगला शब्द अनुमानित करने को क्या उपयोग करता है?
- Images
- तथ्यों का database
- पिछले शब्द के बाद कौन-सा शब्द आया उसकी गिनतियाँ
- Compiler
Answer
पिछले शब्द के बाद कौन-सा शब्द आया उसकी गिनतियाँ — यह शब्द-जोड़ी आवृत्तियों से संभावनाएँ आँकता है।