An evaluation set determines what you believe about your model. Building it carefully matters as much as building the model.
Represent Real Use
Draw examples from actual or realistic inputs, in proportions reflecting real traffic, plus deliberate coverage of important rare cases.
Include Hard and Edge Cases
Ambiguous inputs, unusual formats, adversarial examples and out-of-scope requests.
Quality Labels
- Clear labelling guidelines.
- Expert labellers for specialised domains.
- Multiple labellers and agreement checks for subjective tasks.
Keep It Separate
Never train or tune prompts directly on the test set. Keep a development set for iteration and a held-out test set for final checks.
Size
Large enough to detect meaningful differences; small enough to review. Hundreds of well-chosen examples can be more useful than thousands of random ones.
Slice
Tag examples by category, language, customer segment or difficulty to report performance per slice.
Maintain
Add new failure cases from production, retire outdated examples and version the set.