Securing AI systems: the threat landscape
What actually changes when a model joins a system, the attacks that follow from it, and how to threat model an AI feature before you ship it.
How injection works, why filtering fails, and the design patterns that actually contain it.
Prompt injection is the defining vulnerability of language model applications, and the one most commonly answered with defences that do not work.
This course covers the mechanism, the indirect variants that reach systems through documents and web pages, the defences worth deploying, and how to handle model output without inheriting an older class of bug.
4 lessons · 59 min
One stream of text, no separation of instruction from data, and what follows.
Following an attack from a planted document to a tool call nobody authorised.
Ranking the common mitigations by how much they actually change the system's permissions.
The model is an untrusted string source; treat its output like any other user input.
Sign in and enrol to leave a review.
No reviews yet — be the first once you have worked through it.
1 min read
1 min read
1 min read
1 min read
What actually changes when a model joins a system, the attacks that follow from it, and how to threat model an AI feature before you ship it.
Access control across a retrieval index, the confused deputy problem in tool use, and keeping an agent inside its blast radius.
Attacks on models themselves: evasion, poisoning and backdoors, model theft, and what the data remembers.