पाठ 5 / 28

जाँचा जा सकने वाला सीखा Embedding

TF-IDF और SVD से छोटा embedding बनाएँ और उससे खोजें।

Latent semantic analysis, पुराना विचार छोटे रूप में

बड़े मॉडल के बिना embeddings का काम देखने के लिए हम latent semantic analysis (LSA) उपयोग करते हैं: हर दस्तावेज़ के शब्द-भार का TF-IDF matrix बनाएँ, फिर singular value decomposition (SVD) से उसे कुछ आयामों में संपीड़ित करें। जो शब्द समान दस्तावेज़ों में आते हैं वे समान आयामों में योगदान देते हैं, इसलिए query सिर्फ़ साझा शब्दों से नहीं बल्कि संबंधित शब्दों से भी दस्तावेज़ से मेल खा सकती है। यह असली, सीखा हुआ embedding है, पर छोटा: 12 वाक्य और 6 आयाम, इसलिए मोटा है। आधुनिक neural encoders कहीं अधिक डेटा पर प्रशिक्षित हैं और कहीं ज़्यादा समझते हैं, पर workflow वही है: दस्तावेज़ एक बार embed करें, query embed करें, और vectors की तुलना करें।

सीखें, जाँचें, चुनें

छोटा सीखा embedding प्रक्रिया दिखाता है; असली परियोजनाएँ भाषा, domain, आकार और लागत से मॉडल चुनती हैं।

चार चरण: सीखें, खोजें, चुनें, तैयार करें।
चित्र 2.1 — सीखें, खोजें, चुनें और तैयार करें।

12 वाक्यों पर semantic search, चलाकर

मैंने यह Python virtual environment में numpy 2.5.3, scikit-learn 1.9.1 और faiss-cpu 1.15.1 के साथ चलाया, निश्चित random seeds के साथ ताकि संख्याएँ दोहराई जाएँ। "vacation days" carry-over और annual-leave वाक्य, "stay on a business trip" taxi और hotel वाक्य, और "lost my phone" lost-laptop वाक्य खोजता है। Scores सिर्फ़ इन 12 वाक्यों से सीखे 6-आयामी space में cosine similarities हैं, इसलिए मोटे हैं, और सिर्फ़ 3 आयामों के साथ scores लगभग सब 1.0 थे (कोई विभेद नहीं)।

DOCS = [
 "Employees get 24 days of annual leave and may take a vacation any time of the year",
 "Unused vacation days and leave can be carried over to the next year",
 "Public holidays are paid days off and do not count as annual leave",
 "Sick leave needs a doctor note after three days of absence",
 "Flights must be booked fourteen days in advance for business travel",
 "Hotel stays on a business trip are capped at six thousand rupees per night",
 "Travel expenses need receipts and a trip report within thirty days",
 "Taxi and train tickets are reimbursed for a business trip",
 "Use a password manager and enable two factor authentication on every account",
 "Report a lost laptop or phone to the security team within twenty four hours",
 "Never share a password and always lock your laptop screen",
 "Phishing emails should be reported to the security team immediately",
]
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import TruncatedSVD
from sklearn.preprocessing import normalize

tfidf = TfidfVectorizer(stop_words="english")
X = tfidf.fit_transform(DOCS)
svd = TruncatedSVD(n_components=6, random_state=0)       # learn 6 "topic" dimensions from word co-occurrence
E = normalize(svd.fit_transform(X))                       # one 6-d unit vector per document
print("document embedding shape:", E.shape)

def embed(text): return normalize(svd.transform(tfidf.transform([text])))[0]
def search(q, k=3):
    scores = E @ embed(q)
    return [(round(float(scores[i]), 2), DOCS[i][:48]) for i in np.argsort(-scores)[:k]]

for q in ("how many vacation days do I get", "where can I stay on a business trip", "lost my phone"):
    print(q); [print("   ", r) for r in search(q)]

Output:

document embedding shape: (12, 6)
how many vacation days do I get
    (0.98, 'Unused vacation days and leave can be carried ov')
    (0.97, 'Employees get 24 days of annual leave and may ta')
    (0.82, 'Public holidays are paid days off and do not cou')
where can I stay on a business trip
    (0.96, 'Taxi and train tickets are reimbursed for a busi')
    (0.92, 'Hotel stays on a business trip are capped at six')
    (0.53, 'Flights must be booked fourteen days in advance ')
lost my phone
    (1.0, 'Report a lost laptop or phone to the security te')
    (0.88, 'Phishing emails should be reported to the securi')
    (0.35, 'Never share a password and always lock your lapt')

आयाम एक knob हैं

बहुत कम आयाम सब कुछ धुँधला कर देते हैं; ज़्यादा आयाम विषयों को बेहतर अलग करते हैं जब तक आप शोर को fit न करने लगें। सही आकार डेटा और मॉडल पर निर्भर है।

त्वरित जाँच: LSA "vacation" को "leave" लिखे दस्तावेज़ से कैसे मिलाता है?

  • पर्यायवाची शब्दकोश
  • संबंधित शब्द समान दस्तावेज़ों में आते हैं, इसलिए आयाम साझा करते हैं
  • सटीक string मिलान
  • संयोग
Answer

संबंधित शब्द समान दस्तावेज़ों में आते हैं, इसलिए आयाम साझा करते हैं — Embeddings साथ-साथ आने के पैटर्न से संबंध सीखते हैं।