Lesson 1 of 4
How prompt injection actually works
One stream of text, no separation of instruction from data, and what follows.
15 min 3-question quiz 3 guides to read next
A language model receives a single sequence of tokens. Your system prompt, the user's message and anything you retrieved are concatenated into that sequence, perhaps with role markers. Those markers are a convention the model has been trained to respect, not a boundary it enforces. Sufficiently persuasive text in a lower-trust position can override instructions in a higher-trust one.
That is the whole mechanism. Everything else is technique.
Why filtering does not solve it
The intuitive defence is to detect malicious instructions and strip them. It fails for reasons that are structural rather than temporary:
- The input space is unbounded. "Ignore previous instructions" is trivially blocked and trivially rephrased — in another language, as a story, as base64, as a poem, spread across a document, or as an instruction that reads as perfectly reasonable in isolation.
- Malicious and legitimate text are not distinguishable. "Summarise this and email it to the address at the bottom" is an attack in one context and the feature in another.
- The classifier is a model too. Anything you put in front of the model to screen input can itself be talked around.
Filters raise the cost of an attack. They do not change what the system permits, which is the thing that decides how bad the day gets.
The useful mental model
Think of the model as an enthusiastic contractor who will do anything written on any piece of paper handed to them, cannot tell which papers came from you, and has whatever keys you gave them. You would not fix that by asking them to be more discerning. You would take the keys away, require a signature for anything expensive, and make sure the paperwork they can act on comes from a channel you control.
That is the shape of every defence that works: constrain what the system can do, not what the text can say.
Direct versus indirect
Direct injection comes from the person talking to the model. It matters when the model has authority the user does not — refunds, other users' records, internal data.
Indirect injection comes from content the system reads on the user's behalf: a web page, a PDF, an email, a repository file, an API response. Here the attacker never talks to your system at all; they just leave something where it will be read. This is the variant that scales, and the next lesson follows one end to end.
Check your understanding
3 questions · pass with 2 correct
Enrol for free to save your progress, unlock every lesson and earn a certificate.
Sign in to enrolFurther reading
Guides that go deeper on this lesson.
-
Jailbreaks: How They Work and How to Defend
How people try to get models to bypass their safety training, common techniques, and layered defences.
1 min read
-
Introduction to AI Security
What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.
1 min read
-
Protecting System Prompts
Why system prompts leak, what shouldn't be in them, and how to reduce the impact of extraction.
1 min read