Lesson 28 / 28
Revision: Cheat Sheet and Self-Check
Review the key ideas of the whole course.
Cheat sheet
What: a vector database = durable storage + ANN index + payload filters + updates + operations; families: library, extension (pgvector), dedicated engine (Qdrant, Milvus, Weaviate, Chroma), managed service. Model: collection of records (id, vector, payload); same dimension and metric per collection; deterministic IDs. pgvector: vector(n), operators <-> L2, <=> cosine distance, <#> negative inner product; ORDER BY ... LIMIT k; match the operator class; SQL filters, joins, transactions, RLS, halfvec. Indexes: exact scan first (ground truth); HNSW (m, ef_construction, ef_search) vs IVFFlat (lists, probes); index can be larger than the vectors; measure recall@k vs exact. Filtering: post-filtering can return too few rows (we saw 0); iterative scans, partial indexes or partitions, filter-aware engines with payload indexes; hybrid search with RRF. Scale: batch idempotent ingestion; shard by hashed key (resharding moves most data); replicate, W + R > N; capacity = vectors + links + payload, times copies, plus headroom. Operate: tenant isolation in the database, TLS and encryption, least privilege; backups and tested restores; new model = new index with shadow comparison and rollback; monitor recall, latency percentiles, freshness, cost. Choose: start simple and exact; compare on your data at equal recall; honest benchmarks.
Quick check: You add a `WHERE tenant = 7` filter to an HNSW query and get zero rows although tenant 7 has data. What is the likely cause and a fix?
- The data was deleted by the index
- Post-filtering of a small candidate list; enable iterative scans or use partial indexes/filter-aware search
- Vectors cannot be filtered
- The server clock is wrong
Answer
Post-filtering of a small candidate list; enable iterative scans or use partial indexes/filter-aware search — The index returns nearest candidates first; the filter then removes them.
Quick check: Which statement about HNSW tuning is correct?
- ef_search has no effect
- Raising ef_search raises recall and latency; measure recall against exact search
- Lower ef_search always improves recall
- Recall cannot be measured
Answer
Raising ef_search raises recall and latency; measure recall against exact search — Recall and latency trade off along the ef_search knob.
Quick check: Your cluster keeps 3 copies of each shard. Which write and read settings guarantee you read your latest acknowledged write?
- Neither matters
- Write to 1 and read from 1
- Write to 1 and read from 2 only if lucky
- Write to 2 and read from 2
Answer
Write to 2 and read from 2 — W + R > N guarantees the write and read sets overlap.