Lesson 12 / 28

Choosing an Index

Match index type to data size, update pattern and recall needs.

Size and change rate decide

A practical guide: up to about 100,000 vectors, use flat/brute force, since it is exact and simple. From hundreds of thousands to tens of millions, HNSW is a strong default when memory allows, with IVF variants when memory is tight; add scalar or product quantisation to fit billions or reduce cost. Consider update patterns (frequent inserts and deletes are cheap in some indexes and costly in others), filters (indexes that cannot filter well force you to over-fetch), build time (large indexes can take hours) and whether you need GPU search. Whatever you choose, keep the flat index as a test oracle, record the index parameters with the data version, and re-measure when the data grows.

A rough decision guide

Starting points only; always measure recall and latency yourself.

Vectors            Updates         Memory        Reasonable start
< ~100k             any             any           Flat (exact)
100k - 10M          mostly reads    plenty        HNSW
100k - 10M          mostly reads    tight         IVF (+ scalar quantisation)
10M - billions      batch rebuilds  must shrink   IVF + product quantisation, sharding
heavy filters       any             any           an engine with filter-aware search; over-fetch + post-filter otherwise

Keep the flat index as a test oracle

Even in production, keep a flat index (or a sample of it) to measure the recall of the approximate one over time.

Quick check: Which index is a sensible start for 20,000 vectors?

  • Product quantisation with 32x compression
  • A flat (exact) index
  • A sharded cluster
  • No index at all, guessing
Answer

A flat (exact) index — At small sizes exact search is fast enough and removes tuning.