# Embeddings and Semantic Search — Retrieval-Augmented Generation (RAG)

Source: https://www.geekswithgeeks.com/en/rag/v-embed

> Use vectors to match meaning rather than exact words.

## Closeness means similar meaning

An **embedding model** turns text into a vector of hundreds or thousands of numbers so that texts with similar meaning land close together. To search, embed every chunk once (offline), embed the question at query time and return the chunks whose vectors have the highest **cosine similarity** (or dot product) to it. This finds "carry over unused leave" even if the document says "leftover days can roll into next year". Use the **same embedding model** for chunks and questions, re-embed everything if you change the model, and pick one that supports your languages, since Hindi, English and mixed text may behave differently.

## Find the right passages

Keyword scores match words; embeddings match meaning; hybrid search combines both.

![Five tools: keywords, BM25, vectors, fusion, index.](assets/figures/rag/section-3-map.svg) — Figure 3.1 — Keywords, BM25, vectors, fusion and index.

## A toy embedding, run

I ran this plain-Python (standard library only) example. This is not a trained model: it hashes each word into one of 16 buckets, so it only reflects shared words, but it shows the mechanics of embedding and cosine scoring. The query about unused leave scores 0.722 against the leave sentence and 0.183 against the hotel sentence. Real embedding models also match synonyms.

```python
import hashlib, math

def embed(text, dim=16):
    v = [0.0] * dim
    for w in text.lower().split():
        h = int(hashlib.md5(w.encode()).hexdigest(), 16)
        v[h % dim] += 1 if (h >> 8) % 2 else -1
    n = math.sqrt(sum(x * x for x in v)) or 1
    return [x / n for x in v]

def cos(a, b): return sum(x * y for x, y in zip(a, b))
q = embed("carry over unused leave")
for t in ("unused leave can be carried over", "hotel cost cap per night"):
    print(round(cos(q, embed(t)), 3), t)

```

Output:

```
0.722 unused leave can be carried over
0.183 hotel cost cap per night
```

## Check your languages

If your documents mix Hindi and English, test the embedding model on real mixed queries; multilingual models differ a lot in quality.

**Quiz:** Why must chunks and questions use the same embedding model?

- [ ] It is a legal rule
- [ ] It makes the text shorter
- [x] Vectors from different models are not comparable
- [ ] It removes metadata

*Answer:* Vectors from different models are not comparable. Each model defines its own vector space.
