Lesson 16 / 28

Multi-Vector, Multilingual and Multimodal Embeddings

Know the extensions: many vectors per item, many languages, many media.

One space for more than English text

Multilingual embedding models place sentences from many languages in one shared space, so an English query can retrieve a Hindi document and the reverse; test Hindi, English and mixed Hinglish, as quality varies by language and script. Multimodal models (such as CLIP-style models) embed images and text together, enabling "search images by text" and similar-image search. Multi-vector (late-interaction) methods such as ColBERT keep one vector per token and compare sets of vectors, often more accurate but needing more storage and specialised indexes. A parent-child design embeds small chunks for precise matching but returns the larger parent section. Pick the extension that matches your problem; do not add complexity before the single-vector baseline is measured.

Test Hinglish queries

Users type Hindi in Latin letters and mix languages. Include such queries in your golden set.

Quick check: What does a multilingual embedding model enable?

  • Removing the need for chunking
  • Translating text word for word
  • Matching a query in one language to documents in another
  • Encrypting text
Answer

Matching a query in one language to documents in another — A shared space lets meaning match across languages.