# Model Changes, Re-Indexing and Drift — Embeddings & Vector Search

Source: https://www.geekswithgeeks.com/en/embeddings/e-drift

> Handle upgrades of the embedding model without breaking search.

## A new model means a new space

Vectors from different models, or even different versions of one model, live in **different spaces**. Mixing them silently returns nonsense. The example below makes the point with a rotated copy of a space: distances inside each space are perfectly consistent, so searching "model B" with a "model B" query works, but comparing a "model A" query with "model B" documents returns an unrelated result. Therefore: store the **model name and version** with every vector; when you change models, **re-embed the whole corpus** into a **new index** while the old one keeps serving, compare both on the golden set, then switch over; and keep the ability to roll back. Also watch for **data drift**: new vocabulary, products and languages over time can reduce quality, so re-run the evaluation set on a schedule.

## Mixing two embedding spaces, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. "Model B" is the same data rotated by a random orthogonal matrix. Within model A and within model B the correct document (7) is found. Mixing a model-A query with model-B documents returns document 68, which is meaningless.

```python
import numpy as np
rng = np.random.default_rng(0)
docs = rng.normal(size=(300, 16)); docs /= np.linalg.norm(docs, axis=1, keepdims=True)
q = docs[7] + 0.05 * rng.normal(size=16); q /= np.linalg.norm(q)

Q, _ = np.linalg.qr(rng.normal(size=(16, 16)))                 # "model B" = a rotated copy of the same space
docs_b, q_b = docs @ Q, q @ Q
print("model A, query A vs docs A :", int(np.argmax(docs @ q)), "(correct: 7)")
print("model B, query B vs docs B :", int(np.argmax(docs_b @ q_b)), "(rotation keeps distances: still 7)")
print("MIXED, query A vs docs B   :", int(np.argmax(docs_b @ q)), "(meaningless)")

```

Output:

```
model A, query A vs docs A : 7 (correct: 7)
model B, query B vs docs B : 7 (rotation keeps distances: still 7)
MIXED, query A vs docs B   : 68 (meaningless)
```

## Store the model name with each vector

A model-version field in the metadata makes mixed-model mistakes detectable and re-indexing auditable.

**Quiz:** How should you roll out a new embedding model?

- [ ] Never change models
- [ ] Mix old and new vectors in one index
- [ ] Replace the model and keep the old vectors
- [x] Build a new index alongside the old one, compare, then switch

*Answer:* Build a new index alongside the old one, compare, then switch. A parallel index lets you verify quality and roll back safely.
