Skip to content

Data Poisoning Attacks

How attackers corrupt training or fine-tuning data to change model behaviour, and how to protect data pipelines.

Editorial team 1 min read

Data poisoning manipulates the data a model learns from, so the model behaves as the attacker wants.

Types

  • Availability attacks: degrade overall performance.
  • Targeted attacks: cause specific errors, such as misclassifying one company's products.
  • Backdoors: make the model behave normally except when a trigger appears.

Where Poison Enters

  • Web-scraped training data, where attackers can publish content.
  • User feedback and crowd-sourced labels.
  • Third-party datasets.
  • Documents added to fine-tuning or retrieval corpora.

Research suggests a relatively small number of poisoned documents can be enough to implant some behaviours, even in large training sets.

Defences

  • Know your data's provenance; prefer trusted sources.
  • Validate and filter data: detect duplicates, anomalies and suspicious patterns.
  • Control who can contribute data and labels.
  • Version datasets so changes are traceable.
  • Evaluate models for unexpected behaviour, including trigger testing.
  • Monitor deployed models for behaviour shifts.

RAG Poisoning

Retrieval systems can be poisoned too: planting documents that will be retrieved and influence answers. Control what enters knowledge bases.

More in AI security

All AI security guides →
AI security Guide · 1 min

Introduction to AI Security

What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.

AI security 1 min read 29 Jun 2025

AI security Guide · 1 min

The OWASP Top 10 for LLM Applications

An overview of the widely used list of the most critical security risks for applications built on language models.

AI security 1 min read 28 Jun 2025

AI security Guide · 1 min

Jailbreaks: How They Work and How to Defend

How people try to get models to bypass their safety training, common techniques, and layered defences.

AI security 1 min read 27 Jun 2025

AI security Guide · 1 min

Indirect Prompt Injection

How attackers hide instructions in web pages, emails and documents that AI systems read, and why it's so dangerous for agents.

AI security 1 min read 26 Jun 2025