# Where Embeddings Come From — Embeddings & Vector Search

Source: https://www.geekswithgeeks.com/en/embeddings/b-origin

> See how models learn vectors from data.

## Learned from context

Embeddings are **learned**, not hand-written. Old methods such as **LSA** factorised word-document count matrices; **word2vec** and **GloVe** learned word vectors from the words that appear near each other ("you shall know a word by the company it keeps"). Modern text embeddings come from **transformer encoders** trained with **contrastive learning**: the model sees pairs that should match (a question and its answer, a title and its article, two paraphrases) and pairs that should not, and is trained to pull matching pairs together and push others apart, over huge datasets. The result is a model that maps any text, in many languages, to a vector in a shared space. Different models produce **different spaces**: vectors from two models cannot be compared with each other.

## Training signal in one picture

How contrastive training shapes the space.

```text
positive pair   ("How many leave days?", "Employees get 24 days of annual leave")  -> pull vectors together
negative pair   ("How many leave days?", "Hotels are capped at 6000 rupees")            -> push vectors apart

after training on millions of such pairs:
   [leave questions] [leave policies]  <- one neighbourhood
   [hotel questions] [travel policies] <- another neighbourhood
```

## Check the model card

It lists languages, maximum input length, whether vectors are normalised and any query/document prefixes. Read it before you index anything.

**Quiz:** Why can vectors from two different embedding models not be compared?

- [x] Each model defines its own space and coordinates
- [ ] They always have different lengths in bytes
- [ ] Models forbid it by licence
- [ ] It is only slow

*Answer:* Each model defines its own space and coordinates. Dimension 17 means something different in every model.
