Lesson 22 / 28
Scaling, Cost and Latency
Plan memory, throughput and spending.
Memory is the first wall
Estimate before you build: memory is roughly vectors × dimensions × bytes per number, plus index overhead (graph links for HNSW can add a large fraction) and replicas. When it does not fit: shrink with quantisation, use fewer dimensions, keep full vectors on disk and only compressed ones in memory, shard across machines, or reduce the number of chunks. For latency: cache the embeddings of frequent queries, batch embedding calls, keep the embedding model close to the search service, and measure p95 and p99, not only the average. Costs come from three places: embedding (per-token API fees or GPU time, paid at indexing and for every query), storage/memory, and search compute. Track all three.
Cache query embeddings
Popular queries repeat. Caching their vectors (and even their results) saves embedding fees and latency.
Quick check: What is a first thing to do when vectors do not fit in memory?
- Switch off search
- Ignore the problem
- Delete the metadata only
- Apply quantisation or reduce dimensions, then consider sharding
Answer
Apply quantisation or reduce dimensions, then consider sharding — Compression is the cheapest lever before adding machines.