Lesson 28 / 28
Revision: Cheat Sheet and Self-Check
Review the key ideas of the whole course.
Cheat sheet
Idea: an embedding is a learned vector; similar meaning means nearby vectors; each model has its own space (never mix models). Similarity: dot grows with length and angle, cosine is direction only, L2 is distance; for unit vectors cosine = dot and L2 gives the same ranking, so normalise. Dimensions: more is not automatically better; storage = vectors × dims × bytes; random data concentrates distances. Creating: learned from context (LSA, word2vec) and contrastive training of transformers; choose by languages, domain, max length, dims, cost, licence and your own recall; chunk well, respect max length, use query/document prefixes when required. Search: exact kNN first (ground truth); IVF (nprobe) and HNSW (efSearch) trade recall for speed; scalar and product quantisation trade accuracy for memory. Beyond: metadata filters inside the query, hybrid BM25 + vector with RRF, cross-encoder reranking, clustering, near-duplicate thresholds, multilingual and multimodal models. Quality: golden set with recall@k, MRR, nDCG; error analysis by cause; re-index with a parallel index on model change. Production: library vs database extension vs dedicated engine; incremental updates and deletes; memory arithmetic and p95/p99; vectors as sensitive as source; permissions in the query.
Quick check: A user searches for "time off work" but the policy says "annual leave". Which approach most directly fixes this?
- Exact string matching only
- Embedding (semantic) search, ideally combined with keyword search
- Increasing nprobe
- Using int8 vectors
Answer
Embedding (semantic) search, ideally combined with keyword search — Semantic embeddings bridge vocabulary mismatch; keyword search still covers exact terms.
Quick check: You upgrade the embedding model and only re-embed new documents, leaving the old vectors. What goes wrong?
- Nothing, vectors are interchangeable
- Old and new vectors live in different spaces, so comparisons become meaningless
- The index becomes smaller
- Only latency changes
Answer
Old and new vectors live in different spaces, so comparisons become meaningless — Re-embed everything into a new index when the model changes.
Quick check: Your HNSW index has recall@10 of 0.93 against exact search. What is a sensible next step?
- Switch to a smaller model without testing
- Declare it perfect
- Delete the index
- Raise efSearch (or index parameters) and re-measure against the target recall and latency
Answer
Raise efSearch (or index parameters) and re-measure against the target recall and latency — Recall and latency trade off; tune until the target is met at acceptable cost.