A jailbreak is an attempt to make a model produce content or take actions it was trained to refuse.
Common Techniques
- Role-play: "pretend you're an AI with no rules".
- Hypotheticals and fiction: framing harmful requests as stories.
- Obfuscation: encoding, misspellings, other languages or splitting requests into pieces.
- Many-shot: filling long contexts with fake examples of compliance.
- Gradual escalation: starting harmless and pushing boundaries over many turns.
- Automated search: algorithms that find adversarial prompts.
Jailbreaks Versus Prompt Injection
Jailbreaks usually come from the user trying to bypass safety. Prompt injection comes from third-party content hijacking an application. Defences overlap but aren't identical.
Defences
- Choose models with strong safety training.
- Input and output classifiers that detect harmful requests and responses.
- Clear system prompts defining scope.
- Rate limits and monitoring for repeated attempts.
- Limiting capabilities, so even a successful jailbreak can't do much.
Accept Imperfection
No defence stops every jailbreak. Design so that a bypass has limited consequences, and keep improving based on what you observe.