Datasheets for datasets is a documentation practice, inspired by electronics datasheets, that records the key facts about a dataset for everyone who uses it.
Why Document Datasets
Without documentation, people misuse data: training on it for purposes it doesn't suit, missing gaps in coverage, or breaching its licence. A datasheet makes the dataset's strengths and limits explicit.
Sections to Cover
- Motivation: why was the dataset created, and by whom?
- Composition: what does each instance represent? How many are there? What's missing? Does it contain personal or sensitive data?
- Collection process: how, when and from where was data gathered? Was consent obtained?
- Preprocessing and labelling: what cleaning and labelling was done, by whom, under what guidelines?
- Uses: what has it been used for, and what should it not be used for?
- Distribution: how is it shared, and under what licence?
- Maintenance: who maintains it, how are errors reported, and will it be updated?
Write It as You Build
Documentation is far easier to write while creating the dataset than to reconstruct afterwards.
Keep It Alongside the Data
Store the datasheet with the dataset and version it. A dataset card on a data hub serves the same purpose.
Read Others' Datasheets
When using external data, look for this information. If a dataset has no documentation, treat its suitability with caution.