Data augmentation creates new training examples by applying transformations to existing ones. It helps models generalise, especially with limited data.
Images
Flips, rotations, crops, scaling, colour and brightness changes, blur and noise. Choose transformations that reflect real variation — flipping a photo of a cat is fine; flipping text or medical images with meaningful orientation may not be.
Text
Synonym replacement, back-translation (translate to another language and back), paraphrasing with a language model, and random word deletion. Check that meaning and labels are preserved.
Audio
Background noise, speed and pitch changes, time shifts, and simulated room acoustics.
Tabular Data
Harder to augment safely. Options include adding small noise to numeric features and synthetic oversampling for rare classes (such as SMOTE), applied only to training data.
Guidelines
- Augment training data only; keep validation and test data realistic.
- Keep labels correct after transformation.
- Mirror real-world variation rather than inventing unrealistic examples.
- Measure the effect: compare validation performance with and without augmentation.
Augmentation Is Not a Substitute
It broadens conditions for existing examples but can't create genuinely new information, such as a group or situation absent from the data. Collecting real, representative data remains essential.