Skip to content

Data for Evaluation Sets

How to build test sets that give trustworthy measurements of AI system quality.

Editorial team 1 min read

An evaluation set determines what you believe about your model. Building it carefully matters as much as building the model.

Represent Real Use

Draw examples from actual or realistic inputs, in proportions reflecting real traffic, plus deliberate coverage of important rare cases.

Include Hard and Edge Cases

Ambiguous inputs, unusual formats, adversarial examples and out-of-scope requests.

Quality Labels

  • Clear labelling guidelines.
  • Expert labellers for specialised domains.
  • Multiple labellers and agreement checks for subjective tasks.

Keep It Separate

Never train or tune prompts directly on the test set. Keep a development set for iteration and a held-out test set for final checks.

Size

Large enough to detect meaningful differences; small enough to review. Hundreds of well-chosen examples can be more useful than thousands of random ones.

Slice

Tag examples by category, language, customer segment or difficulty to report performance per slice.

Maintain

Add new failure cases from production, retire outdated examples and version the set.

More in Data for AI

All Data for AI guides →
Data for AI Guide · 1 min

How to Read a Dataset Card

The questions a dataset card should answer — what one row is, where the data came from, its licence and its quirks — before you use it.

Data for AI 1 min read 5 Apr 2026

Data for AI Guide · 2 min

Data Quality Dimensions

Accuracy, completeness, consistency, timeliness, validity and uniqueness: a framework for checking whether data is fit for purpose.

Data for AI 2 min read 4 Apr 2026

Data for AI Guide · 2 min

Handling Missing Data

Why data goes missing, how to find out, and the options — dropping, imputing, flagging — with their trade-offs.

Data for AI 2 min read 3 Apr 2026

Data for AI Guide · 2 min

Detecting and Handling Outliers

How to spot unusual values, decide whether they are errors or genuine extremes, and treat them appropriately.

Data for AI 2 min read 2 Apr 2026