Skip to content

Data Versioning

Why datasets need versions like code does, and practical approaches from snapshots to dedicated tools.

Editorial team 2 min read

Models and reports depend on data that changes. Without versioning, you can't reproduce a result or explain why it changed.

Why Version Data

  • Reproducibility: retrain a model or rerun an analysis on exactly the same data.
  • Debugging: find out whether a change in results came from code or data.
  • Auditing: show what data a decision or model was based on.
  • Collaboration: everyone works from the same known version.

Approaches

  • Dated snapshots: save extracts with dates in their names and keep them read-only. Simple and effective.
  • Content hashes: record a hash of each file so you can prove which version was used.
  • Dataset versions from publishers: record the version or revision identifier of external datasets.
  • Table formats with time travel: some lakehouse table formats keep history so you can query data as it was at a past point.
  • Data version control tools: track large files alongside code in version control.

What to Record

For every model or significant analysis: the dataset name, version or snapshot, date extracted, filters applied and the code version that processed it.

Version Splits and Labels Too

Train, validation and test splits and labelling guideline versions affect results as much as the raw data.

Balance Cost

Keeping every version of huge datasets is expensive. Keep versions used by production models and published results, and define retention for the rest.

More in Data for AI

All Data for AI guides →
Data for AI Guide · 1 min

How to Read a Dataset Card

The questions a dataset card should answer — what one row is, where the data came from, its licence and its quirks — before you use it.

Data for AI 1 min read 5 Apr 2026

Data for AI Guide · 2 min

Data Quality Dimensions

Accuracy, completeness, consistency, timeliness, validity and uniqueness: a framework for checking whether data is fit for purpose.

Data for AI 2 min read 4 Apr 2026

Data for AI Guide · 2 min

Handling Missing Data

Why data goes missing, how to find out, and the options — dropping, imputing, flagging — with their trade-offs.

Data for AI 2 min read 3 Apr 2026

Data for AI Guide · 2 min

Detecting and Handling Outliers

How to spot unusual values, decide whether they are errors or genuine extremes, and treat them appropriately.

Data for AI 2 min read 2 Apr 2026