Skip to content

Jailbreaks: How They Work and How to Defend

How people try to get models to bypass their safety training, common techniques, and layered defences.

Editorial team 1 min read

A jailbreak is an attempt to make a model produce content or take actions it was trained to refuse.

Common Techniques

  • Role-play: "pretend you're an AI with no rules".
  • Hypotheticals and fiction: framing harmful requests as stories.
  • Obfuscation: encoding, misspellings, other languages or splitting requests into pieces.
  • Many-shot: filling long contexts with fake examples of compliance.
  • Gradual escalation: starting harmless and pushing boundaries over many turns.
  • Automated search: algorithms that find adversarial prompts.

Jailbreaks Versus Prompt Injection

Jailbreaks usually come from the user trying to bypass safety. Prompt injection comes from third-party content hijacking an application. Defences overlap but aren't identical.

Defences

  • Choose models with strong safety training.
  • Input and output classifiers that detect harmful requests and responses.
  • Clear system prompts defining scope.
  • Rate limits and monitoring for repeated attempts.
  • Limiting capabilities, so even a successful jailbreak can't do much.

Accept Imperfection

No defence stops every jailbreak. Design so that a bypass has limited consequences, and keep improving based on what you observe.

More in AI security

All AI security guides →
AI security Guide · 1 min

Introduction to AI Security

What AI security covers — attacks on models, data and AI applications — and how it differs from traditional security.

AI security 1 min read 29 Jun 2025

AI security Guide · 1 min

The OWASP Top 10 for LLM Applications

An overview of the widely used list of the most critical security risks for applications built on language models.

AI security 1 min read 28 Jun 2025

AI security Guide · 1 min

Indirect Prompt Injection

How attackers hide instructions in web pages, emails and documents that AI systems read, and why it's so dangerous for agents.

AI security 1 min read 26 Jun 2025

AI security Guide · 1 min

Data Poisoning Attacks

How attackers corrupt training or fine-tuning data to change model behaviour, and how to protect data pipelines.

AI security 1 min read 25 Jun 2025