Duplicates distort counts, bias models, and cause models to memorise repeated content.
Exact Duplicates
Hash each record — or a normalised version of it — and group identical hashes. Fast and simple.
Near-Duplicates
Records that differ slightly: extra whitespace, updated timestamps, minor edits.
Techniques:
- Normalisation: lowercasing, trimming, standardising formats before comparing.
- MinHash and locality-sensitive hashing: efficiently find documents with high overlap.
- Embeddings: find semantically similar items.
- Fuzzy matching: string similarity for names and addresses.
Entity Resolution
Deciding whether records refer to the same real-world entity — the same customer across systems — combines matching rules, similarity scores and sometimes machine learning.
For AI Training Data
Deduplicating training data reduces memorisation, improves generalisation and prevents test data leaking into training.
Choosing What to Keep
Define rules: the most recent, most complete or most trusted source.
Document
Record how deduplication was performed and how many records were removed.