Skip to content

Choosing an Embedding Model for RAG

The factors that matter when picking an embedding model for retrieval: language, domain, dimensions, input length, cost and licence.

Editorial team 2 min read

The embedding model decides which passages look similar to a question. Choosing well has a large effect on retrieval quality.

Factors to Weigh

  • Languages: multilingual models are needed if documents or users span languages.
  • Domain: general models may underperform on legal, medical or code content; domain-specific models exist.
  • Maximum input length: chunks longer than the limit are truncated.
  • Vector dimensions: larger vectors can capture more but cost more to store and search. Some models support shorter vectors with modest quality loss.
  • Speed and cost: embedding millions of chunks, and every query, adds up.
  • Hosting and licence: API-based or self-hosted open models; check licence terms.

Use Benchmarks to Shortlist

Public retrieval benchmarks compare models across tasks, but your content may differ. Treat leaderboard rankings as a starting point.

Test on Your Data

Embed your documents with two or three candidates and measure recall@k on your real questions. The best model on a leaderboard is not always best for your documents.

Asymmetric Models

Some models expect different prefixes or modes for queries and documents. Follow the model card's instructions exactly, or quality drops.

Switching Models

Changing models means re-embedding everything, because vectors from different models aren't comparable. Choose carefully and plan migrations.

Fine-Tuning Embeddings

With enough question–passage pairs from your domain, fine-tuning an embedding model can improve retrieval noticeably.

More in RAG

All RAG guides →
RAG Guide · 1 min

RAG Architecture: The Components End to End

A map of a complete retrieval-augmented generation system, from ingestion to answer, and what each component is responsible for.

RAG 1 min read 6 Dec 2025

RAG Guide · 2 min

Document Parsing for RAG

Turning PDFs, slides, HTML and scans into clean, structured text — the unglamorous step that decides RAG quality.

RAG 2 min read 5 Dec 2025

RAG Guide · 2 min

Chunk Size and Overlap Tuning

How to choose chunk size and overlap for retrieval by testing against real questions rather than guessing.

RAG 2 min read 4 Dec 2025

RAG Guide · 2 min

Hybrid Search for RAG

Combining keyword and vector search so RAG finds both exact terms and paraphrased meaning.

RAG 2 min read 3 Dec 2025