The data used to train AI raises questions about consent, copyright and fairness to the people and creators behind it.
Sources of Training Data
- Data you collect directly from users or customers.
- Licensed datasets.
- Public datasets with published licences.
- Web-scraped content.
- Synthetic data.
Personal Data
If training data contains personal information, privacy law applies: you need a lawful basis, must respect purpose limits, and may need to handle deletion requests. Data collected for one purpose can't automatically be used to train models for another.
Copyright and Licences
Datasets and content come with licences that may restrict commercial use, require attribution or prohibit redistribution. The legal position on training AI with copyrighted material is contested and varies by country. Record the source and licence of every dataset you use.
Respecting Creators
Consider whether creators expected their work to be used this way, honour opt-outs where offered (such as robots.txt and other machine-readable signals), and prefer licensed or openly licensed content.
Good Practice
- Maintain a data inventory with sources, licences and consent basis.
- Prefer datasets with clear documentation and licences.
- Remove personal data you don't need.
- Provide ways for people to opt out where appropriate.
- Get legal advice for significant commercial uses.
Why It Matters
Beyond legal risk, respecting data rights builds trust with users, customers and creators.