Lesson 1 of 4
Evasion: adversarial examples
Small, deliberate changes to an input that flip a model's answer.
15 min 3-question quiz 3 guides to read next
An adversarial example is an input modified so slightly that a person sees no difference, yet the model's prediction changes completely. The canonical demonstration is an image of a panda that a classifier labels a gibbon with high confidence after adding noise invisible to the eye.
Why it happens
A model carves its input space into regions by decision boundary. In high dimensions those boundaries pass surprisingly close to ordinary inputs, so a small step in the right direction crosses one. The direction is not mysterious: if you have the model, you can compute the gradient of its loss with respect to the input and step along it. That is the fast gradient sign method, and every stronger attack is a refinement of the idea.
The uncomfortable properties
They transfer. An example crafted against one model often fools another trained on similar data, even with a different architecture. So an attacker does not need your weights — a surrogate will do, which makes "our model is private" a weak defence.
Black-box attacks work. With only query access, an attacker can estimate gradients from the outputs, or simply search. Confidence scores make this much faster, which is one reason to return labels rather than probabilities on a public endpoint.
They exist in the physical world. Printed patterns on a sign, a sticker on a lens, a pattern on a shirt. Robustness to camera angle and lighting is achievable by averaging the attack over those transformations.
Where it actually matters
Not everywhere. The risk is real when a model gates something an adversary wants: content moderation, fraud scoring, malware classification, biometric access, spam filtering. In those settings assume a motivated adversary who will probe, adapt and keep what works.
For a model predicting next quarter's demand, the adversarial threat is mostly theoretical — and the real risk is the data drifting, which the next lessons are closer to.
Check your understanding
3 questions · pass with 2 correct
Enrol for free to save your progress, unlock every lesson and earn a certificate.
Sign in to enrolFurther reading
Guides that go deeper on this lesson.
-
Adversarial Examples
How tiny, carefully crafted changes to inputs can fool machine learning models, and approaches to robustness.
1 min read
-
AI-Powered Cyber Attacks
How attackers use AI to scale phishing, find vulnerabilities and automate attacks, and what it means for defenders.
1 min read
-
Introduction to AI Security
What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.
1 min read