Lesson 28 / 28

Revision: Cheat Sheet and Self-Check

Review the key ideas of the whole course.

Cheat sheet

Idea: an embedding is a learned vector; similar meaning means nearby vectors; each model has its own space (never mix models). Similarity: dot grows with length and angle, cosine is direction only, L2 is distance; for unit vectors cosine = dot and L2 gives the same ranking, so normalise. Dimensions: more is not automatically better; storage = vectors × dims × bytes; random data concentrates distances. Creating: learned from context (LSA, word2vec) and contrastive training of transformers; choose by languages, domain, max length, dims, cost, licence and your own recall; chunk well, respect max length, use query/document prefixes when required. Search: exact kNN first (ground truth); IVF (nprobe) and HNSW (efSearch) trade recall for speed; scalar and product quantisation trade accuracy for memory. Beyond: metadata filters inside the query, hybrid BM25 + vector with RRF, cross-encoder reranking, clustering, near-duplicate thresholds, multilingual and multimodal models. Quality: golden set with recall@k, MRR, nDCG; error analysis by cause; re-index with a parallel index on model change. Production: library vs database extension vs dedicated engine; incremental updates and deletes; memory arithmetic and p95/p99; vectors as sensitive as source; permissions in the query.

Quick check: A user searches for "time off work" but the policy says "annual leave". Which approach most directly fixes this?

  • Exact string matching only
  • Embedding (semantic) search, ideally combined with keyword search
  • Increasing nprobe
  • Using int8 vectors
Answer

Embedding (semantic) search, ideally combined with keyword search — Semantic embeddings bridge vocabulary mismatch; keyword search still covers exact terms.

Quick check: You upgrade the embedding model and only re-embed new documents, leaving the old vectors. What goes wrong?

  • Nothing, vectors are interchangeable
  • Old and new vectors live in different spaces, so comparisons become meaningless
  • The index becomes smaller
  • Only latency changes
Answer

Old and new vectors live in different spaces, so comparisons become meaningless — Re-embed everything into a new index when the model changes.

Quick check: Your HNSW index has recall@10 of 0.93 against exact search. What is a sensible next step?

  • Switch to a smaller model without testing
  • Declare it perfect
  • Delete the index
  • Raise efSearch (or index parameters) and re-measure against the target recall and latency
Answer

Raise efSearch (or index parameters) and re-measure against the target recall and latency — Recall and latency trade off; tune until the target is met at acceptable cost.