Lesson 5 / 28
A Learned Embedding You Can Inspect
Build a small embedding with TF-IDF and SVD and search with it.
Latent semantic analysis, the old idea in miniature
To see embeddings work without a large model, we use latent semantic analysis (LSA): build a TF-IDF matrix of word weights per document, then use singular value decomposition (SVD) to compress it into a few dimensions. Words that appear in similar documents end up contributing to the same dimensions, so a query can match a document through related words, not only shared words. It is a real, learned embedding, but tiny: 12 sentences and 6 dimensions, so it is crude. Modern neural encoders are trained on vastly more data and understand far more, but the workflow is identical: embed the documents once, embed the query, and compare vectors.
Learn them, inspect them, choose them
A tiny learned embedding shows the mechanics; real projects choose a model by language, domain, size and cost.
Semantic search on 12 sentences, run
I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. "vacation days" finds the carry-over and annual-leave sentences, "stay on a business trip" finds the taxi and hotel sentences, and "lost my phone" finds the lost-laptop sentence. The scores are cosine similarities in a 6-dimension space learned from only these 12 sentences, so they are crude, and with only 3 dimensions the scores were nearly all 1.0 (no discrimination).
DOCS = [
"Employees get 24 days of annual leave and may take a vacation any time of the year",
"Unused vacation days and leave can be carried over to the next year",
"Public holidays are paid days off and do not count as annual leave",
"Sick leave needs a doctor note after three days of absence",
"Flights must be booked fourteen days in advance for business travel",
"Hotel stays on a business trip are capped at six thousand rupees per night",
"Travel expenses need receipts and a trip report within thirty days",
"Taxi and train tickets are reimbursed for a business trip",
"Use a password manager and enable two factor authentication on every account",
"Report a lost laptop or phone to the security team within twenty four hours",
"Never share a password and always lock your laptop screen",
"Phishing emails should be reported to the security team immediately",
]
import numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import TruncatedSVD
from sklearn.preprocessing import normalize
tfidf = TfidfVectorizer(stop_words="english")
X = tfidf.fit_transform(DOCS)
svd = TruncatedSVD(n_components=6, random_state=0) # learn 6 "topic" dimensions from word co-occurrence
E = normalize(svd.fit_transform(X)) # one 6-d unit vector per document
print("document embedding shape:", E.shape)
def embed(text): return normalize(svd.transform(tfidf.transform([text])))[0]
def search(q, k=3):
scores = E @ embed(q)
return [(round(float(scores[i]), 2), DOCS[i][:48]) for i in np.argsort(-scores)[:k]]
for q in ("how many vacation days do I get", "where can I stay on a business trip", "lost my phone"):
print(q); [print(" ", r) for r in search(q)]
Output:
document embedding shape: (12, 6)
how many vacation days do I get
(0.98, 'Unused vacation days and leave can be carried ov')
(0.97, 'Employees get 24 days of annual leave and may ta')
(0.82, 'Public holidays are paid days off and do not cou')
where can I stay on a business trip
(0.96, 'Taxi and train tickets are reimbursed for a busi')
(0.92, 'Hotel stays on a business trip are capped at six')
(0.53, 'Flights must be booked fourteen days in advance ')
lost my phone
(1.0, 'Report a lost laptop or phone to the security te')
(0.88, 'Phishing emails should be reported to the securi')
(0.35, 'Never share a password and always lock your lapt')Dimensions are a knob
Too few dimensions blur everything together; more dimensions separate topics better until you start fitting noise. The right size depends on data and model.
Quick check: What lets LSA match "vacation" with a document that says "leave"?
- A dictionary of synonyms
- Related words appear in the same documents, so they share dimensions
- Exact string matching
- Random chance
Answer
Related words appear in the same documents, so they share dimensions — Embeddings learn relatedness from patterns of co-occurrence.