Jailbreak
A crafted prompt that tricks a model into ignoring its safety rules and producing content it is meant to refuse.
It is a core safety concern and the reason guardrails need constant testing.
Frequently asked questions
What is an AI jailbreak?
It is a crafted prompt that tricks a model into ignoring its safety rules and producing content it is meant to refuse. Think of it as talking the AI out of its own guardrails through clever wording.
How is a jailbreak different from prompt injection?
A jailbreak is usually the user directly persuading the model to break its rules, while prompt injection hides malicious instructions inside outside content the AI reads. In a jailbreak the person chatting is the one pushing the limits.
Why do AI companies keep testing for jailbreaks?
Because guardrails need constant testing to stay effective as people invent new tricks. Jailbreaks are a core safety concern, so finding them early is how providers patch weaknesses before they cause harm.
New to all this? Start with what an AI agent really is, browse the full glossary, or explore the learning hub.
← Back to the glossary