Skip to content

Lesson 1 of 4

Free preview

Why AI systems break differently

Instructions and data stop being separable, and the system stops being deterministic.

14 min 3-question quiz 3 guides to read next

On this page
  1. What is actually new
  2. What is not new at all

Most of security rests on two assumptions that a language model quietly breaks.

The first is that instructions and data are different things. A database distinguishes a query from the values in it; that is what parameterised queries are for. A language model has no such separation. It receives one stream of text, and anything in that stream can read as an instruction — including the contents of a document it was asked to summarise. This is why prompt injection is not a bug to be patched but a property of the component.

The second is that the same input produces the same output. Testing, signatures and most detection logic assume this. A model's output varies with temperature, with the surrounding context, and across versions of the model itself. A payload that is blocked today may succeed tomorrow with a slightly different phrasing, and a test that passed yesterday proves less than you would like.

What is actually new

Three things genuinely change:

  • The parser is the attacker's playground. Natural language is an enormous input space with no grammar you can validate against. "Ignore previous instructions" is the naive version; the effective versions are indirect, encoded or buried in a file the user never read.
  • The system acts. A model that only writes text is a content problem. A model wired to tools — sending email, running queries, calling APIs — is an execution problem, and the question becomes what it can reach, not what it can say.
  • The data is the attack surface. Training sets, fine-tuning data, retrieval corpora and context windows all become places to put something malicious, and none of them look like code.

What is not new at all

Most incidents involving AI systems are ordinary failures wearing a new hat: an over-permissioned API key in a prompt template, a vector database exposed to the internet, a tool endpoint with no authorisation check, logs full of customer data shipped to a third party. The model makes these worse because it will happily be talked into exercising them, but the underlying mistakes are the ones we have made for twenty years.

The practical consequence: treat the model as an untrusted, persuadable component sitting inside your trust boundary. Everything it can reach, an attacker can eventually reach through it. The rest of this course is about drawing that boundary deliberately rather than by accident.

Check your understanding

3 questions · pass with 2 correct

1. Which long-standing assumption does a language model break?
2. Why is prompt injection described as a property rather than a bug?
3. What is the practical way to treat a model inside a system?

You'll see your score; enrol to have it count towards your certificate.

Enrol for free to save your progress, unlock every lesson and earn a certificate.

Sign in to enrol

Further reading

Guides that go deeper on this lesson.