RAG systems fail in two places: retrieval and generation. Evaluate both, separately and together.
Build a Test Set
Collect realistic questions — ideally from real users — each with:
- the expected answer or key facts;
- the source passages that contain it;
- questions that shouldn't be answered (out of scope or not in the documents).
Retrieval Metrics
- Recall@k: does the right passage appear in the top k?
- Precision@k / MRR: how much noise, and how high the first relevant result ranks.
Answer Metrics
- Correctness: does the answer match the expected facts?
- Faithfulness (groundedness): is every claim supported by the retrieved context?
- Answer relevance: does it address the question?
- Citation quality: are citations present and supportive?
- Refusal accuracy: does it decline when the information isn't there?
How to Grade
Use exact checks where possible, model-based graders with clear rubrics for faithfulness and relevance, and human review on a sample to validate the graders.
Diagnose by Stage
If retrieval recall is low, fix parsing, chunking, search or filters. If retrieval is good but answers are poor, fix the prompt, context assembly or model.
Run It Continuously
Re-evaluate whenever documents, chunking, embedding models, prompts or language models change, and track results over time.