Pipelines often carry personal data from source systems into warehouses, analytics and AI. Protecting it is a core engineering responsibility.
Classify
Tag columns and datasets containing personal or sensitive data, ideally automatically, and record them in catalogues.
Minimise
Only copy personal fields that downstream uses need.
Protect
- Masking: hide values from users who don't need them.
- Tokenisation and pseudonymisation: replace identifiers with tokens, keeping joins possible.
- Encryption in transit and at rest.
- Row- and column-level access policies in warehouses.
Non-Production Environments
Don't copy real personal data into development and testing environments; use masked or synthetic data.
Deletion Requests
Design pipelines so records can be deleted across all copies when people exercise their rights.
Logging
Avoid personal data in pipeline logs and error messages.
Audit
Track who accesses sensitive data and review regularly.