Jailbreak
Güvenlik kısıtlarını aşma (Jailbreak)
D5 · Security, privacy, governance and intellectual property
In language-model security, a jailbreak is an attempt to bypass safeguards so a model produces otherwise restricted behavior.
Review status: 2026-12-05
Technical explanation
The target is a safety restriction; a successful bypass is evaluated against the system’s intended restrictions.
Conceptual boundaries
Usage overlaps with prompt injection; OWASP treats jailbreaking as a form of it. This is not device operating-system jailbreaking.
Provider-neutral example
Illustrative: an authorized evaluator checks whether a test assistant maintains a defined safety boundary.
Limitations
Results depend on the tested model and safeguards; a successful test does not establish a universal bypass.
Related concepts
Atomic claims and evidence
1.1OWASP LLM01:2025 Prompt Injection
- Source
- OWASP LLM01:2025 Prompt Injection
- Source role
- Authoritative source
- Exact locator
- Opening definition: relationship between prompt injection and jailbreaking; Prevention and Mitigation Strategies
- Supported claim
- OWASP describes jailbreaking as bypassing safety protocols; Zou and colleagues study circumventing alignment safeguards.
- Last verification
- Review due
- Scope limitation
- This is a conceptual security definition, not an attack procedure or assurance about current products.
1.2Universal and Transferable Adversarial Attacks on Aligned Language Models
- Source
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Source role
- Authoritative source
- Exact locator
- Abstract: circumventing alignment measures and objectionable generation
- Supported claim
- OWASP describes jailbreaking as bypassing safety protocols; Zou and colleagues study circumventing alignment safeguards.
- Last verification
- Review due
- Scope limitation
- This is a conceptual security definition, not an attack procedure or assurance about current products.