Lesson 6 / 28

Keyword Search vs Semantic Search

See where exact-word matching falls short and where it still wins.

Complementary, not competing

Keyword search (TF-IDF, BM25) matches the words that appear. It is excellent for exact terms: product codes, error numbers, names, rare jargon. It fails on vocabulary mismatch: a query for "time off work" shares almost no words with a document about "annual leave". Semantic (embedding) search bridges that gap by comparing meaning, but it can be fuzzy on exact identifiers and rare terms, and it always returns something nearest, even when nothing is relevant. Real systems usually combine both (hybrid search) and add a score threshold so that irrelevant queries return "nothing found".

Keyword matching on unseen wording, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. "time off work" has no word in common with "annual leave" sentences; it still scores 0.408 against the sentence that happens to contain the word "time", which is luck, not understanding. The single word "vacation" matches because that exact word appears in a sentence.

DOCS = [
 "Employees get 24 days of annual leave and may take a vacation any time of the year",
 "Unused vacation days and leave can be carried over to the next year",
 "Public holidays are paid days off and do not count as annual leave",
 "Sick leave needs a doctor note after three days of absence",
 "Flights must be booked fourteen days in advance for business travel",
 "Hotel stays on a business trip are capped at six thousand rupees per night",
 "Travel expenses need receipts and a trip report within thirty days",
 "Taxi and train tickets are reimbursed for a business trip",
 "Use a password manager and enable two factor authentication on every account",
 "Report a lost laptop or phone to the security team within twenty four hours",
 "Never share a password and always lock your laptop screen",
 "Phishing emails should be reported to the security team immediately",
]
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

tfidf = TfidfVectorizer(stop_words="english")
X = tfidf.fit_transform(DOCS)
for q in ("time off work", "vacation"):
    s = cosine_similarity(tfidf.transform([q]), X)[0]
    i = int(np.argmax(s))
    print(f"{q!r:16} best keyword match: score {s[i]:.3f} ->", DOCS[i][:50])

Output:

'time off work'  best keyword match: score 0.408 -> Employees get 24 days of annual leave and may take
'vacation'       best keyword match: score 0.416 -> Unused vacation days and leave can be carried over

Always add a "no match" threshold

Nearest-neighbour search always returns something. A minimum score lets the system say it found nothing relevant.

Quick check: When does keyword search beat embeddings?

  • Finding synonyms
  • Matching paraphrases
  • Cross-language matching
  • Looking up an exact error code or product id
Answer

Looking up an exact error code or product id — Exact identifiers have no "meaning" to generalise; the literal token matters.