Lesson 4 / 28

Dimensions and the Curse of Dimensionality

Understand why distances behave oddly in high dimensions.

Everything becomes equally far

More dimensions let a model store more nuance, but also cost more memory and search time (storage is vectors × dimensions × bytes per number). In very high dimensions raw distances also concentrate: the nearest and farthest points to a query become almost equally far, so naive distance becomes less discriminating for random data. Real embeddings are not random, since meaningful data sits on lower-dimensional structure, which is why 384 to 1536 dimensions work well in practice. Practical lessons: more dimensions is not automatically better; test smaller models or dimension-reduced vectors; some modern models support truncating to fewer dimensions with modest quality loss (check the documentation); and always measure retrieval quality on your own data.

Distance concentration, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. For 2,000 random points, the ratio of nearest to farthest distance rises from 0.004 in 2 dimensions to 0.965 in 10,000 dimensions: in high dimensions random points are nearly all equally far. This is a property of random data, not of trained embeddings.

import numpy as np
rng = np.random.default_rng(0)
print("dims | nearest vs farthest distance ratio (1.0 means everything is equally far)")
for d in (2, 10, 100, 1000, 10000):
    X = rng.uniform(size=(2000, d)); q = rng.uniform(size=d)
    dist = np.linalg.norm(X - q, axis=1)
    print(f"{d:5d} | {dist.min() / dist.max():.3f}")

Output:

dims | nearest vs farthest distance ratio (1.0 means everything is equally far)
    2 | 0.004
   10 | 0.253
  100 | 0.665
 1000 | 0.894
10000 | 0.965

Storage arithmetic, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. One million 768-dimension float32 vectors take 3.1 GB; one hundred million take 307 GB. Halving the dimensions halves the size, and storing numbers as float16 or int8 shrinks it by 2x or 4x.

def gb(n_vectors, dims, bytes_per): return n_vectors * dims * bytes_per / 1e9

for n in (1_000_000, 100_000_000):
    print(f"{n:>11,} vectors x 768 dims: float32 {gb(n,768,4):8.1f} GB | float16 {gb(n,768,2):8.1f} GB | int8 {gb(n,768,1):7.1f} GB | 384 dims float32 {gb(n,384,4):8.1f} GB")

Output:

  1,000,000 vectors x 768 dims: float32      3.1 GB | float16      1.5 GB | int8     0.8 GB | 384 dims float32      1.5 GB
100,000,000 vectors x 768 dims: float32    307.2 GB | float16    153.6 GB | int8    76.8 GB | 384 dims float32    153.6 GB

Truncating dimensions: test first

Only models trained for it keep quality when you keep the first N dimensions. Measure recall before and after.

Quick check: Is a higher-dimensional embedding model always better?

  • Yes, always
  • No: it costs more storage and time, so test on your data
  • Only if it is free
  • Dimensions never matter
Answer

No: it costs more storage and time, so test on your data — Quality gains from extra dimensions often flatten while costs keep growing.