The embedding model decides which passages look similar to a question. Choosing well has a large effect on retrieval quality.
Factors to Weigh
- Languages: multilingual models are needed if documents or users span languages.
- Domain: general models may underperform on legal, medical or code content; domain-specific models exist.
- Maximum input length: chunks longer than the limit are truncated.
- Vector dimensions: larger vectors can capture more but cost more to store and search. Some models support shorter vectors with modest quality loss.
- Speed and cost: embedding millions of chunks, and every query, adds up.
- Hosting and licence: API-based or self-hosted open models; check licence terms.
Use Benchmarks to Shortlist
Public retrieval benchmarks compare models across tasks, but your content may differ. Treat leaderboard rankings as a starting point.
Test on Your Data
Embed your documents with two or three candidates and measure recall@k on your real questions. The best model on a leaderboard is not always best for your documents.
Asymmetric Models
Some models expect different prefixes or modes for queries and documents. Follow the model card's instructions exactly, or quality drops.
Switching Models
Changing models means re-embedding everything, because vectors from different models aren't comparable. Choose carefully and plan migrations.
Fine-Tuning Embeddings
With enough question–passage pairs from your domain, fine-tuning an embedding model can improve retrieval noticeably.