पाठ 6 / 28
Keyword Search बनाम Semantic Search
देखें सटीक-शब्द मिलान कहाँ कम पड़ता है और कहाँ अब भी जीतता है।
प्रतिस्पर्धी नहीं, पूरक
Keyword search (TF-IDF, BM25) आए हुए शब्दों से मिलान करती है। यह सटीक शब्दों के लिए उत्कृष्ट है: उत्पाद कोड, error संख्याएँ, नाम, दुर्लभ jargon। यह शब्दावली-बेमेल पर विफल होती है: "time off work" की query "annual leave" वाले दस्तावेज़ से लगभग कोई शब्द साझा नहीं करती। Semantic (embedding) search अर्थ की तुलना करके यह अंतर पाटती है, पर सटीक पहचानकर्ताओं और दुर्लभ शब्दों पर धुँधली हो सकती है, और हमेशा कुछ निकटतम लौटाती है, भले कुछ प्रासंगिक न हो। असली systems आम तौर पर दोनों मिलाते हैं (hybrid search) और score सीमा जोड़ते हैं ताकि अप्रासंगिक queries "कुछ नहीं मिला" लौटाएँ।
अनदेखे शब्दांकन पर keyword matching, चलाकर
मैंने यह Python virtual environment में numpy 2.5.3, scikit-learn 1.9.1 और faiss-cpu 1.15.1 के साथ चलाया, निश्चित random seeds के साथ ताकि संख्याएँ दोहराई जाएँ। "time off work" का "annual leave" वाक्यों से कोई साझा शब्द नहीं; फिर भी उस वाक्य से 0.408 score पाता है जिसमें संयोग से शब्द "time" है, जो भाग्य है, समझ नहीं। एकल शब्द "vacation" इसलिए मिलता है क्योंकि वही शब्द एक वाक्य में है।
DOCS = [
"Employees get 24 days of annual leave and may take a vacation any time of the year",
"Unused vacation days and leave can be carried over to the next year",
"Public holidays are paid days off and do not count as annual leave",
"Sick leave needs a doctor note after three days of absence",
"Flights must be booked fourteen days in advance for business travel",
"Hotel stays on a business trip are capped at six thousand rupees per night",
"Travel expenses need receipts and a trip report within thirty days",
"Taxi and train tickets are reimbursed for a business trip",
"Use a password manager and enable two factor authentication on every account",
"Report a lost laptop or phone to the security team within twenty four hours",
"Never share a password and always lock your laptop screen",
"Phishing emails should be reported to the security team immediately",
]
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np
tfidf = TfidfVectorizer(stop_words="english")
X = tfidf.fit_transform(DOCS)
for q in ("time off work", "vacation"):
s = cosine_similarity(tfidf.transform([q]), X)[0]
i = int(np.argmax(s))
print(f"{q!r:16} best keyword match: score {s[i]:.3f} ->", DOCS[i][:50])
Output:
'time off work' best keyword match: score 0.408 -> Employees get 24 days of annual leave and may take 'vacation' best keyword match: score 0.416 -> Unused vacation days and leave can be carried over
हमेशा "कोई मेल नहीं" सीमा जोड़ें
Nearest-neighbour खोज हमेशा कुछ लौटाती है। न्यूनतम score system को कहने देता है कि कुछ प्रासंगिक नहीं मिला।
त्वरित जाँच: Keyword search embeddings को कब हराती है?
- पर्यायवाची खोजना
- पुनर्कथन मिलाना
- क्रॉस-भाषा मिलान
- सटीक error code या product id खोजना
Answer
सटीक error code या product id खोजना — सटीक पहचानकर्ताओं में सामान्यीकरण का कोई "अर्थ" नहीं; शाब्दिक token मायने रखता है।