Data poisoning manipulates the data a model learns from, so the model behaves as the attacker wants.
Types
- Availability attacks: degrade overall performance.
- Targeted attacks: cause specific errors, such as misclassifying one company's products.
- Backdoors: make the model behave normally except when a trigger appears.
Where Poison Enters
- Web-scraped training data, where attackers can publish content.
- User feedback and crowd-sourced labels.
- Third-party datasets.
- Documents added to fine-tuning or retrieval corpora.
Research suggests a relatively small number of poisoned documents can be enough to implant some behaviours, even in large training sets.
Defences
- Know your data's provenance; prefer trusted sources.
- Validate and filter data: detect duplicates, anomalies and suspicious patterns.
- Control who can contribute data and labels.
- Version datasets so changes are traceable.
- Evaluate models for unexpected behaviour, including trigger testing.
- Monitor deployed models for behaviour shifts.
RAG Poisoning
Retrieval systems can be poisoned too: planting documents that will be retrieved and influence answers. Control what enters knowledge bases.