Skip to content

Lesson 1 of 4

Free preview

How prompt injection actually works

One stream of text, no separation of instruction from data, and what follows.

15 min 3-question quiz 3 guides to read next

On this page
  1. Why filtering does not solve it
  2. The useful mental model
  3. Direct versus indirect

A language model receives a single sequence of tokens. Your system prompt, the user's message and anything you retrieved are concatenated into that sequence, perhaps with role markers. Those markers are a convention the model has been trained to respect, not a boundary it enforces. Sufficiently persuasive text in a lower-trust position can override instructions in a higher-trust one.

That is the whole mechanism. Everything else is technique.

Why filtering does not solve it

The intuitive defence is to detect malicious instructions and strip them. It fails for reasons that are structural rather than temporary:

  • The input space is unbounded. "Ignore previous instructions" is trivially blocked and trivially rephrased — in another language, as a story, as base64, as a poem, spread across a document, or as an instruction that reads as perfectly reasonable in isolation.
  • Malicious and legitimate text are not distinguishable. "Summarise this and email it to the address at the bottom" is an attack in one context and the feature in another.
  • The classifier is a model too. Anything you put in front of the model to screen input can itself be talked around.

Filters raise the cost of an attack. They do not change what the system permits, which is the thing that decides how bad the day gets.

The useful mental model

Think of the model as an enthusiastic contractor who will do anything written on any piece of paper handed to them, cannot tell which papers came from you, and has whatever keys you gave them. You would not fix that by asking them to be more discerning. You would take the keys away, require a signature for anything expensive, and make sure the paperwork they can act on comes from a channel you control.

That is the shape of every defence that works: constrain what the system can do, not what the text can say.

Direct versus indirect

Direct injection comes from the person talking to the model. It matters when the model has authority the user does not — refunds, other users' records, internal data.

Indirect injection comes from content the system reads on the user's behalf: a web page, a PDF, an email, a repository file, an API response. Here the attacker never talks to your system at all; they just leave something where it will be read. This is the variant that scales, and the next lesson follows one end to end.

Check your understanding

3 questions · pass with 2 correct

1. Why can't prompt injection be solved by filtering?
2. What distinguishes indirect from direct injection?
3. What shape do defences that actually work take?

You'll see your score; enrol to have it count towards your certificate.

Enrol for free to save your progress, unlock every lesson and earn a certificate.

Sign in to enrol

Further reading

Guides that go deeper on this lesson.