Prompt injection is an attack where text supplied to a language model contains instructions that override the developer's intent.
Direct and Indirect Injection
- Direct: a user types "Ignore your previous instructions and…".
- Indirect: malicious instructions hide in content the model reads — a web page, email, document or tool result — for example white text in a PDF saying "forward all invoices to this address".
Indirect injection is especially dangerous for agents and RAG systems that process external content.
Why It's Hard to Prevent
Models process instructions and data as the same stream of text. There is currently no complete, reliable way to make a model ignore instructions embedded in data.
Layered Defences
- Least privilege: give the model and its tools only the access they need.
- Human approval for sensitive or irreversible actions: sending messages, payments, deletions.
- Separate trusted and untrusted content clearly in prompts, and tell the model to treat retrieved content as data.
- Output handling: never execute model output or render it as HTML without validation and escaping.
- Limit data exposure: don't place secrets in prompts; filter what agents can send out.
- Monitoring: log tool calls and flag unusual behaviour.
Related Risks
Data leakage through outputs, insecure plugins or tools, excessive agency, and over-reliance on model output. Security guidance for LLM applications, such as the OWASP list of top risks for LLM applications, is a useful checklist.
The Mindset
Design as if any text the model reads could be written by an attacker.