Skip to content

Lesson 1 of 4

Free preview

Evasion: adversarial examples

Small, deliberate changes to an input that flip a model's answer.

15 min 3-question quiz 3 guides to read next

On this page
  1. Why it happens
  2. The uncomfortable properties
  3. Where it actually matters

An adversarial example is an input modified so slightly that a person sees no difference, yet the model's prediction changes completely. The canonical demonstration is an image of a panda that a classifier labels a gibbon with high confidence after adding noise invisible to the eye.

Why it happens

A model carves its input space into regions by decision boundary. In high dimensions those boundaries pass surprisingly close to ordinary inputs, so a small step in the right direction crosses one. The direction is not mysterious: if you have the model, you can compute the gradient of its loss with respect to the input and step along it. That is the fast gradient sign method, and every stronger attack is a refinement of the idea.

The uncomfortable properties

They transfer. An example crafted against one model often fools another trained on similar data, even with a different architecture. So an attacker does not need your weights — a surrogate will do, which makes "our model is private" a weak defence.

Black-box attacks work. With only query access, an attacker can estimate gradients from the outputs, or simply search. Confidence scores make this much faster, which is one reason to return labels rather than probabilities on a public endpoint.

They exist in the physical world. Printed patterns on a sign, a sticker on a lens, a pattern on a shirt. Robustness to camera angle and lighting is achievable by averaging the attack over those transformations.

Where it actually matters

Not everywhere. The risk is real when a model gates something an adversary wants: content moderation, fraud scoring, malware classification, biometric access, spam filtering. In those settings assume a motivated adversary who will probe, adapt and keep what works.

For a model predicting next quarter's demand, the adversarial threat is mostly theoretical — and the real risk is the data drifting, which the next lessons are closer to.

Check your understanding

3 questions · pass with 2 correct

1. What is an adversarial example?
2. Why does transferability matter for defenders?
3. Why return labels rather than full probability vectors on a public endpoint?

You'll see your score; enrol to have it count towards your certificate.

Enrol for free to save your progress, unlock every lesson and earn a certificate.

Sign in to enrol

Further reading

Guides that go deeper on this lesson.

  • Adversarial Examples

    How tiny, carefully crafted changes to inputs can fool machine learning models, and approaches to robustness.

    1 min read

  • AI-Powered Cyber Attacks

    How attackers use AI to scale phishing, find vulnerabilities and automate attacks, and what it means for defenders.

    1 min read

  • Introduction to AI Security

    What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.

    1 min read