# Dimensions and the Curse of Dimensionality — Embeddings & Vector Search

Source: https://www.geekswithgeeks.com/en/embeddings/b-dims

> Understand why distances behave oddly in high dimensions.

## Everything becomes equally far

More dimensions let a model store more nuance, but also cost more memory and search time (storage is `vectors × dimensions × bytes per number`). In very high dimensions raw distances also **concentrate**: the nearest and farthest points to a query become almost equally far, so naive distance becomes less discriminating for random data. Real embeddings are not random, since meaningful data sits on lower-dimensional structure, which is why 384 to 1536 dimensions work well in practice. Practical lessons: more dimensions is **not automatically better**; test smaller models or dimension-reduced vectors; some modern models support **truncating** to fewer dimensions with modest quality loss (check the documentation); and always measure retrieval quality on your own data.

## Distance concentration, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. For 2,000 random points, the ratio of nearest to farthest distance rises from 0.004 in 2 dimensions to 0.965 in 10,000 dimensions: in high dimensions random points are nearly all equally far. This is a property of random data, not of trained embeddings.

```python
import numpy as np
rng = np.random.default_rng(0)
print("dims | nearest vs farthest distance ratio (1.0 means everything is equally far)")
for d in (2, 10, 100, 1000, 10000):
    X = rng.uniform(size=(2000, d)); q = rng.uniform(size=d)
    dist = np.linalg.norm(X - q, axis=1)
    print(f"{d:5d} | {dist.min() / dist.max():.3f}")

```

Output:

```
dims | nearest vs farthest distance ratio (1.0 means everything is equally far)
    2 | 0.004
   10 | 0.253
  100 | 0.665
 1000 | 0.894
10000 | 0.965
```

## Storage arithmetic, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. One million 768-dimension float32 vectors take 3.1 GB; one hundred million take 307 GB. Halving the dimensions halves the size, and storing numbers as float16 or int8 shrinks it by 2x or 4x.

```python
def gb(n_vectors, dims, bytes_per): return n_vectors * dims * bytes_per / 1e9

for n in (1_000_000, 100_000_000):
    print(f"{n:>11,} vectors x 768 dims: float32 {gb(n,768,4):8.1f} GB | float16 {gb(n,768,2):8.1f} GB | int8 {gb(n,768,1):7.1f} GB | 384 dims float32 {gb(n,384,4):8.1f} GB")

```

Output:

```
  1,000,000 vectors x 768 dims: float32      3.1 GB | float16      1.5 GB | int8     0.8 GB | 384 dims float32      1.5 GB
100,000,000 vectors x 768 dims: float32    307.2 GB | float16    153.6 GB | int8    76.8 GB | 384 dims float32    153.6 GB
```

## Truncating dimensions: test first

Only models trained for it keep quality when you keep the first N dimensions. Measure recall before and after.

**Quiz:** Is a higher-dimensional embedding model always better?

- [ ] Yes, always
- [x] No: it costs more storage and time, so test on your data
- [ ] Only if it is free
- [ ] Dimensions never matter

*Answer:* No: it costs more storage and time, so test on your data. Quality gains from extra dimensions often flatten while costs keep growing.
