पाठ 10 / 27

BM25: मानक Keyword Ranker

अधिकतर search engines के पीछे की ranking function उपयोग करें।

Saturation और लंबाई नियंत्रण वाला TF-IDF

BM25 TF-IDF को दो विचारों से सुधारता है: term-frequency saturation (किसी शब्द की दसवीं उपस्थिति दूसरी से बहुत कम जोड़ती है) और लंबाई normalisation (छोटे दस्तावेज़ में मिलान, बहुत लंबे में उसी मिलान से ज़्यादा गिना जाता है)। दो पैरामीटर, आम तौर पर k1 लगभग 1.2 से 2 और b लगभग 0.75, इन्हें नियंत्रित करते हैं। BM25 Elasticsearch, OpenSearch और Lucene में डिफ़ॉल्ट है और मज़बूत baseline है जो सटीक-शब्द queries पर अक्सर बुरी तरह चुने embeddings को हरा देता है। Vectors जोड़ने पर भी इसे उपलब्ध रखें।

BM25 शुरू से, चलाकर

मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। "expense approval director" expense नीति को पहले (2.534) और "lost laptop report" security नीति को पहले (3.024) रखता है; अन्य दस्तावेज़ शून्य के पास score पाते हैं।

DOCS = {
 "leave": "Employees get 24 days of paid leave per year. Unused leave up to 5 days can be carried over to the next year.",
 "remote": "Remote work is allowed up to 3 days per week with manager approval. Core hours are 11:00 to 16:00.",
 "expense": "Expenses above 5000 rupees need approval from a director. Submit receipts within 30 days.",
 "security": "Use a password manager and enable two factor authentication. Report lost laptops within 24 hours.",
 "travel": "Flights must be booked at least 14 days in advance. Hotel cost is capped at 6000 rupees per night.",
}
import math, re
from collections import Counter

def tok(t): return re.findall(r"[a-z0-9]+", t.lower())
names = list(DOCS); docs = [tok(DOCS[n]) for n in names]
N = len(docs); avg = sum(len(d) for d in docs) / N
df = Counter(w for d in docs for w in set(d))

def bm25(q, d, k1=1.5, b=0.75):
    tf = Counter(d); score = 0.0
    for w in tok(q):
        if w not in tf: continue
        idf = math.log((N - df[w] + 0.5) / (df[w] + 0.5) + 1)
        score += idf * tf[w] * (k1 + 1) / (tf[w] + k1 * (1 - b + b * len(d) / avg))
    return score

for q in ("expense approval director", "lost laptop report"):
    ranked = sorted(((round(bm25(q, d), 3), n) for n, d in zip(names, docs)), reverse=True)[:3]
    print(q, "->", ranked)

Output:

expense approval director -> [(2.534, 'expense'), (0.823, 'remote'), (0.0, 'travel')]
lost laptop report -> [(3.024, 'security'), (0.0, 'travel'), (0.0, 'remote')]

Identifiers के लिए BM25 रखें

Order संख्याएँ, error codes और नाम अक्सर embeddings चूक जाते हैं। Keyword index उन्हें भरोसे से पकड़ता है।

त्वरित जाँच: BM25 की लंबाई normalisation क्या करती है?

  • लंबे दस्तावेज़ों को सिर्फ़ ज़्यादा शब्द होने से जीतने से रोकती है
  • Query छोटी करती है
  • पाठ encrypt करती है
  • Embedding model चुनती है
Answer

लंबे दस्तावेज़ों को सिर्फ़ ज़्यादा शब्द होने से जीतने से रोकती है — संक्षिप्त दस्तावेज़ में मिलान विशाल दस्तावेज़ में उसी मिलान से मज़बूत प्रमाण है।