The test set is the foundation of RAG improvement. Without it, every change is guesswork.
Sources of Questions
- Real user questions from logs, search queries or support tickets (with personal data removed).
- Questions written by subject-matter experts, including tricky ones.
- Generated questions from documents, reviewed by people — useful for coverage, but often easier than real questions.
What to Record for Each Question
- The question.
- The expected answer or key facts.
- The documents or passages that contain the answer.
- Whether it should be answered at all.
- Category and difficulty.
Coverage
Include common questions, rare but important ones, multi-part questions, ambiguous questions, questions whose answers changed recently, and questions outside scope.
Size
Start with 50–100 well-chosen questions; expand as you find failures. Keep a stable core for comparisons over time.
Keep It Honest
Don't tune prompts using the exact test questions as examples. Keep a hidden subset for final checks.
Maintain It
When documents change, update expected answers. Add every production failure you fix as a new test question.
Share It
A test set agreed with content owners and stakeholders defines what "good" means for everyone.