Synthetic data is artificially generated data that mimics the statistical properties of real data without being real records.
How It's Generated
- Rule-based: generate values from defined distributions and business rules. Good for software testing.
- Statistical models: learn distributions and correlations from real data and sample new records.
- Deep generative models: learn complex patterns, including in images and text.
- Simulation: model a process, such as traffic or a factory, and record the outputs.
Where It Helps
- Testing and development without exposing production data.
- Sharing data with partners or researchers when real data is sensitive.
- Rare scenarios — edge cases and failures that seldom occur in real data.
- Balancing under-represented classes.
Privacy Isn't Automatic
Synthetic data generated from real records can still leak information, especially about unusual individuals. Assess re-identification risk, and consider techniques with formal guarantees, such as differential privacy, for sensitive data.
Realism Limits
Synthetic data captures the patterns the generator learned — and misses others. Models trained only on synthetic data may perform worse on real data.
Validate It
- Compare distributions and relationships with the real data.
- Train on synthetic, test on real, and compare with training on real data.
- Check privacy risk.
Label It
Mark synthetic data clearly so it isn't mistaken for real data later.