Skip to content

Building Representative Datasets

Why training data must reflect the people and situations a model will face, and how to check and improve coverage.

Editorial team 2 min read

A model learns from its data. If the data doesn't represent the real world it will be used in, the model will fail — often for the people least represented.

Common Gaps

  • Groups: under-representation of certain ages, regions, languages, accents or skin tones.
  • Conditions: images only in good lighting; speech only from quiet rooms.
  • Time: data from a period that no longer reflects current behaviour.
  • Channels: data from one source, such as web users only, when the model will serve phone users too.

Checking Representativeness

  1. Describe the population and conditions the model will face.
  2. Compare the dataset's make-up against that description.
  3. Evaluate model performance separately for important groups and conditions.
  4. Pay special attention to small groups, where performance is often worst.

Improving Coverage

  • Collect more data from under-represented groups and conditions.
  • Partner with communities to collect data respectfully and with consent.
  • Use stratified sampling when building datasets.
  • Use augmentation carefully to broaden conditions.
  • Reweight examples during training.

Document Limitations

When gaps can't be fixed, state them in the dataset and model cards, and restrict use accordingly.

Revisit Over Time

Populations and behaviour change. Periodically compare production data with training data and refresh datasets.

More in Data for AI

All Data for AI guides →
Data for AI Guide · 1 min

How to Read a Dataset Card

The questions a dataset card should answer — what one row is, where the data came from, its licence and its quirks — before you use it.

Data for AI 1 min read 5 Apr 2026

Data for AI Guide · 2 min

Data Quality Dimensions

Accuracy, completeness, consistency, timeliness, validity and uniqueness: a framework for checking whether data is fit for purpose.

Data for AI 2 min read 4 Apr 2026

Data for AI Guide · 2 min

Handling Missing Data

Why data goes missing, how to find out, and the options — dropping, imputing, flagging — with their trade-offs.

Data for AI 2 min read 3 Apr 2026

Data for AI Guide · 2 min

Detecting and Handling Outliers

How to spot unusual values, decide whether they are errors or genuine extremes, and treat them appropriately.

Data for AI 2 min read 2 Apr 2026