Lesson 28 / 28

Revision: Cheat Sheet and Self-Check

Review the key ideas of the whole course.

Cheat sheet

What: a vector database = durable storage + ANN index + payload filters + updates + operations; families: library, extension (pgvector), dedicated engine (Qdrant, Milvus, Weaviate, Chroma), managed service. Model: collection of records (id, vector, payload); same dimension and metric per collection; deterministic IDs. pgvector: vector(n), operators <-> L2, <=> cosine distance, <#> negative inner product; ORDER BY ... LIMIT k; match the operator class; SQL filters, joins, transactions, RLS, halfvec. Indexes: exact scan first (ground truth); HNSW (m, ef_construction, ef_search) vs IVFFlat (lists, probes); index can be larger than the vectors; measure recall@k vs exact. Filtering: post-filtering can return too few rows (we saw 0); iterative scans, partial indexes or partitions, filter-aware engines with payload indexes; hybrid search with RRF. Scale: batch idempotent ingestion; shard by hashed key (resharding moves most data); replicate, W + R > N; capacity = vectors + links + payload, times copies, plus headroom. Operate: tenant isolation in the database, TLS and encryption, least privilege; backups and tested restores; new model = new index with shadow comparison and rollback; monitor recall, latency percentiles, freshness, cost. Choose: start simple and exact; compare on your data at equal recall; honest benchmarks.

Quick check: You add a `WHERE tenant = 7` filter to an HNSW query and get zero rows although tenant 7 has data. What is the likely cause and a fix?

  • The data was deleted by the index
  • Post-filtering of a small candidate list; enable iterative scans or use partial indexes/filter-aware search
  • Vectors cannot be filtered
  • The server clock is wrong
Answer

Post-filtering of a small candidate list; enable iterative scans or use partial indexes/filter-aware search — The index returns nearest candidates first; the filter then removes them.

Quick check: Which statement about HNSW tuning is correct?

  • ef_search has no effect
  • Raising ef_search raises recall and latency; measure recall against exact search
  • Lower ef_search always improves recall
  • Recall cannot be measured
Answer

Raising ef_search raises recall and latency; measure recall against exact search — Recall and latency trade off along the ef_search knob.

Quick check: Your cluster keeps 3 copies of each shard. Which write and read settings guarantee you read your latest acknowledged write?

  • Neither matters
  • Write to 1 and read from 1
  • Write to 1 and read from 2 only if lucky
  • Write to 2 and read from 2
Answer

Write to 2 and read from 2 — W + R > N guarantees the write and read sets overlap.