# Scaling, Cost and Latency — Embeddings & Vector Search

Source: https://www.geekswithgeeks.com/en/embeddings/p-scale

> Plan memory, throughput and spending.

## Memory is the first wall

Estimate before you build: memory is roughly `vectors × dimensions × bytes per number`, **plus index overhead** (graph links for HNSW can add a large fraction) and replicas. When it does not fit: shrink with **quantisation**, use **fewer dimensions**, keep full vectors on disk and only compressed ones in memory, **shard** across machines, or reduce the number of chunks. For latency: cache the embeddings of frequent queries, batch embedding calls, keep the embedding model close to the search service, and measure **p95 and p99**, not only the average. Costs come from three places: **embedding** (per-token API fees or GPU time, paid at indexing and for every query), **storage/memory**, and **search compute**. Track all three.

## Cache query embeddings

Popular queries repeat. Caching their vectors (and even their results) saves embedding fees and latency.

**Quiz:** What is a first thing to do when vectors do not fit in memory?

- [ ] Switch off search
- [ ] Ignore the problem
- [ ] Delete the metadata only
- [x] Apply quantisation or reduce dimensions, then consider sharding

*Answer:* Apply quantisation or reduce dimensions, then consider sharding. Compression is the cheapest lever before adding machines.
