Models and reports depend on data that changes. Without versioning, you can't reproduce a result or explain why it changed.
Why Version Data
- Reproducibility: retrain a model or rerun an analysis on exactly the same data.
- Debugging: find out whether a change in results came from code or data.
- Auditing: show what data a decision or model was based on.
- Collaboration: everyone works from the same known version.
Approaches
- Dated snapshots: save extracts with dates in their names and keep them read-only. Simple and effective.
- Content hashes: record a hash of each file so you can prove which version was used.
- Dataset versions from publishers: record the version or revision identifier of external datasets.
- Table formats with time travel: some lakehouse table formats keep history so you can query data as it was at a past point.
- Data version control tools: track large files alongside code in version control.
What to Record
For every model or significant analysis: the dataset name, version or snapshot, date extracted, filters applied and the code version that processed it.
Version Splits and Labels Too
Train, validation and test splits and labelling guideline versions affect results as much as the raw data.
Balance Cost
Keeping every version of huge datasets is expensive. Keep versions used by production models and published results, and define retention for the rest.