# Embeddings and Similarity — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/n-embeddings

> Represent tokens as vectors where distance reflects meaning.

## Meaning as coordinates

Neural networks work on numbers, so each token is mapped to a list of numbers called an **embedding** (hundreds to thousands of values). During training, tokens used in similar contexts end up with similar vectors. **Cosine similarity** (the angle between two vectors) measures closeness: near 1 means very similar direction, near 0 unrelated. The same idea powers **semantic search**: embed documents and the query, then return the nearest vectors. The numbers below are hand-made to show the idea; real embeddings are learned.

## Numbers in, probabilities out

Embeddings turn tokens into vectors; softmax turns scores into probabilities; sampling picks a token.

![Five stages: embed, score, softmax, sample, repeat.](assets/figures/llms/section-2-map.svg) — Figure 2.1 — Embed, score, softmax, sample and repeat.

## Cosine similarity, run

I ran this plain-Python (standard library only) example. "king" and "queen" point in almost the same direction (0.994); "king" and "apple" do not (0.303).

```python
import math

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))

emb = {
    "king":   [0.9, 0.8, 0.1],
    "queen":  [0.9, 0.7, 0.2],
    "apple":  [0.1, 0.2, 0.9],
}
print("king~queen", round(cosine(emb["king"], emb["queen"]), 3))
print("king~apple", round(cosine(emb["king"], emb["apple"]), 3))

```

Output:

```
king~queen 0.994
king~apple 0.303
```

## Embeddings are model-specific

Vectors from different embedding models are not interchangeable. If you switch models, re-embed all your documents.

**Quiz:** What does a cosine similarity near 1 indicate?

- [ ] The vectors are unrelated
- [x] The vectors point in nearly the same direction
- [ ] The text is long
- [ ] The model is large

*Answer:* The vectors point in nearly the same direction. Similar meaning gives similar direction in embedding space.
