# Multi-Vector, Multilingual and Multimodal Embeddings — Embeddings & Vector Search

Source: https://www.geekswithgeeks.com/en/embeddings/a-multi

> Know the extensions: many vectors per item, many languages, many media.

## One space for more than English text

**Multilingual** embedding models place sentences from many languages in one shared space, so an English query can retrieve a Hindi document and the reverse; test Hindi, English and mixed Hinglish, as quality varies by language and script. **Multimodal** models (such as CLIP-style models) embed images and text together, enabling "search images by text" and similar-image search. **Multi-vector (late-interaction)** methods such as ColBERT keep one vector per token and compare sets of vectors, often more accurate but needing more storage and specialised indexes. A **parent-child** design embeds small chunks for precise matching but returns the larger parent section. Pick the extension that matches your problem; do not add complexity before the single-vector baseline is measured.

## Test Hinglish queries

Users type Hindi in Latin letters and mix languages. Include such queries in your golden set.

**Quiz:** What does a multilingual embedding model enable?

- [ ] Removing the need for chunking
- [ ] Translating text word for word
- [x] Matching a query in one language to documents in another
- [ ] Encrypting text

*Answer:* Matching a query in one language to documents in another. A shared space lets meaning match across languages.
