Lesson 7 / 28
Choosing an Embedding Model
Weigh language, domain, size, cost and licence.
There is no universally best model
Options include hosted APIs (a provider embeds text for a per-token fee; easy, but data leaves your system and you depend on the provider) and open models you run yourself (sentence-transformers-style encoders, multilingual models; you control data and cost, but you manage hardware). Criteria: languages (a model must support Hindi if your users write Hindi, and mixed Hindi-English text needs testing), domain (legal, medical, code), maximum input length (long chunks get truncated), dimensions (storage and speed), latency and throughput, price, licence, and quality on your data. Public leaderboards such as MTEB are a useful shortlist, but they average over tasks that may not resemble yours, so evaluate 2 to 3 candidates on a small labelled set of your own questions before you commit. Record the exact model name and version you used.
A model-selection checklist
Use it to compare candidates fairly.
[ ] supports all my languages (test real Hindi / Hinglish queries)
[ ] max input tokens >= my chunk size (otherwise text is silently cut)
[ ] dimensions x vectors fits my memory/budget (see the storage arithmetic)
[ ] latency / throughput OK for indexing the corpus and for live queries
[ ] licence and data-handling terms acceptable (hosted: where does the text go?)
[ ] recall@k / MRR measured on 50+ of MY questions, for 2-3 candidates
[ ] exact model name + version recorded; plan for a future re-indexPin the model version
Providers retire and update embedding models. Pin a version and plan the re-index in advance.
Quick check: Why test candidate models on your own questions rather than rely on a leaderboard alone?
- Leaderboards average over tasks that may not resemble your domain or languages
- Leaderboards are always wrong
- Your questions are secret
- It is required by law
Answer
Leaderboards average over tasks that may not resemble your domain or languages — Use leaderboards to shortlist and your own data to decide.