Lesson 1 of 4
Why AI systems break differently
Instructions and data stop being separable, and the system stops being deterministic.
14 min 3-question quiz 3 guides to read next
On this page
Most of security rests on two assumptions that a language model quietly breaks.
The first is that instructions and data are different things. A database distinguishes a query from the values in it; that is what parameterised queries are for. A language model has no such separation. It receives one stream of text, and anything in that stream can read as an instruction — including the contents of a document it was asked to summarise. This is why prompt injection is not a bug to be patched but a property of the component.
The second is that the same input produces the same output. Testing, signatures and most detection logic assume this. A model's output varies with temperature, with the surrounding context, and across versions of the model itself. A payload that is blocked today may succeed tomorrow with a slightly different phrasing, and a test that passed yesterday proves less than you would like.
What is actually new
Three things genuinely change:
- The parser is the attacker's playground. Natural language is an enormous input space with no grammar you can validate against. "Ignore previous instructions" is the naive version; the effective versions are indirect, encoded or buried in a file the user never read.
- The system acts. A model that only writes text is a content problem. A model wired to tools — sending email, running queries, calling APIs — is an execution problem, and the question becomes what it can reach, not what it can say.
- The data is the attack surface. Training sets, fine-tuning data, retrieval corpora and context windows all become places to put something malicious, and none of them look like code.
What is not new at all
Most incidents involving AI systems are ordinary failures wearing a new hat: an over-permissioned API key in a prompt template, a vector database exposed to the internet, a tool endpoint with no authorisation check, logs full of customer data shipped to a third party. The model makes these worse because it will happily be talked into exercising them, but the underlying mistakes are the ones we have made for twenty years.
The practical consequence: treat the model as an untrusted, persuadable component sitting inside your trust boundary. Everything it can reach, an attacker can eventually reach through it. The rest of this course is about drawing that boundary deliberately rather than by accident.
Check your understanding
3 questions · pass with 2 correct
Enrol for free to save your progress, unlock every lesson and earn a certificate.
Sign in to enrolFurther reading
Guides that go deeper on this lesson.
-
Introduction to AI Security
What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.
1 min read
-
Threat Modelling AI Applications
A structured way to identify how an AI system could be attacked or misused before building defences.
1 min read
-
AI Security Frameworks and Standards
An overview of frameworks for managing AI security: NIST AI RMF, MITRE ATLAS, ISO/IEC 42001, OWASP and more.
1 min read