Skip to content

Data Cleaning Best Practices

How to clean data reproducibly: fixing types, standardising values, removing duplicates and documenting every step.

Editorial team 2 min read

Data cleaning often takes most of a project's time. Doing it systematically makes results trustworthy and repeatable.

Clean in Code, Not by Hand

Write cleaning steps as scripts or notebooks that run from the raw data. Manual edits in spreadsheets can't be repeated on next month's data and leave no record.

Keep the Raw Data

Never overwrite the original. Clean into a new dataset so you can always start again.

Common Steps

  • Fix types: parse dates, convert numeric text, keep identifiers as strings.
  • Standardise text: trim whitespace, consistent case, unify spellings ("NSW", "N.S.W.", "New South Wales").
  • Handle missing values deliberately.
  • Remove or merge duplicates, defining what counts as the same record.
  • Validate ranges and codes against rules.
  • Standardise units and currencies.
  • Fix structural issues: split combined columns, reshape wide and long tables.

Check After Every Step

Compare row counts and summary statistics before and after each change. Unexpected drops in rows often reveal a bad join or filter.

Document Decisions

Record what you changed and why: which outliers were removed, how duplicates were resolved, what assumptions were made.

Fix Problems at the Source

When the same problem recurs, raise it with the data owner. Fixing data entry or collection is better than cleaning the same mess forever.

Test Your Cleaning

Write simple checks — no duplicate IDs, no negative prices — that run automatically on cleaned data.

More in Data for AI

All Data for AI guides →
Data for AI Guide · 1 min

How to Read a Dataset Card

The questions a dataset card should answer — what one row is, where the data came from, its licence and its quirks — before you use it.

Data for AI 1 min read 5 Apr 2026

Data for AI Guide · 2 min

Data Quality Dimensions

Accuracy, completeness, consistency, timeliness, validity and uniqueness: a framework for checking whether data is fit for purpose.

Data for AI 2 min read 4 Apr 2026

Data for AI Guide · 2 min

Handling Missing Data

Why data goes missing, how to find out, and the options — dropping, imputing, flagging — with their trade-offs.

Data for AI 2 min read 3 Apr 2026

Data for AI Guide · 2 min

Detecting and Handling Outliers

How to spot unusual values, decide whether they are errors or genuine extremes, and treat them appropriately.

Data for AI 2 min read 2 Apr 2026