Lesson 2 / 28

Where Embeddings Come From

See how models learn vectors from data.

Learned from context

Embeddings are learned, not hand-written. Old methods such as LSA factorised word-document count matrices; word2vec and GloVe learned word vectors from the words that appear near each other ("you shall know a word by the company it keeps"). Modern text embeddings come from transformer encoders trained with contrastive learning: the model sees pairs that should match (a question and its answer, a title and its article, two paraphrases) and pairs that should not, and is trained to pull matching pairs together and push others apart, over huge datasets. The result is a model that maps any text, in many languages, to a vector in a shared space. Different models produce different spaces: vectors from two models cannot be compared with each other.

Training signal in one picture

How contrastive training shapes the space.

positive pair   ("How many leave days?", "Employees get 24 days of annual leave")  -> pull vectors together
negative pair   ("How many leave days?", "Hotels are capped at 6000 rupees")            -> push vectors apart

after training on millions of such pairs:
   [leave questions] [leave policies]  <- one neighbourhood
   [hotel questions] [travel policies] <- another neighbourhood

Check the model card

It lists languages, maximum input length, whether vectors are normalised and any query/document prefixes. Read it before you index anything.

Quick check: Why can vectors from two different embedding models not be compared?

  • Each model defines its own space and coordinates
  • They always have different lengths in bytes
  • Models forbid it by licence
  • It is only slow
Answer

Each model defines its own space and coordinates — Dimension 17 means something different in every model.