In search and RAG systems, the answer can only be as good as what's retrieved. Measure retrieval separately from generation.
Build a Retrieval Test Set
Collect realistic queries and, for each, the documents or chunks that should be found. Real user questions are best; subject-matter experts can mark the relevant passages.
Core Metrics
- Recall@k: of the relevant items, what share appears in the top k results? Crucial for RAG, where only the top few chunks reach the model.
- Precision@k: of the top k results, what share are relevant?
- Hit rate@k: did at least one relevant item appear in the top k?
- MRR (mean reciprocal rank): how high the first relevant result appears, on average.
- NDCG: rewards ranking highly relevant items near the top, and supports graded relevance.
Choosing k
Use the k that matches your system: if you pass five chunks to the model, measure recall@5.
Diagnose Failures
For queries where retrieval fails, look at what was returned instead. Common causes: poor chunking, missing metadata, synonyms the index doesn't handle, and outdated content.
Track Over Time
Re-run the retrieval evaluation whenever you change the embedding model, chunking, index settings or document set.
Then Evaluate Answers
Once retrieval is solid, evaluate whether the generated answers are correct and supported by the retrieved passages.