◈ AI GLOSSARY ◈

Prompt Injection

An attack where hidden instructions in a document, webpage, or message trick an AI into ignoring its real rules.

WHY IT MATTERS

It is a top security risk for any agent that reads outside content. Untrusted text can hijack behavior.

RELATED TERMS

Frequently asked questions

Can a prompt injection steal my data?

It can, if your agent has tools that touch sensitive data and it gets tricked into misusing them. Hidden instructions in a document or webpage can tell the AI to leak or send information, which is why agents that read outside content need tight guardrails.

How is prompt injection different from a jailbreak?

A jailbreak is when a user directly tricks the model into breaking its own safety rules, while prompt injection hides malicious instructions inside content the AI reads, like a webpage or email. In injection, the attacker may not even be the person chatting with the AI.

How do I protect an AI agent from prompt injection?

Treat any outside text the agent reads as untrusted, limit what tools and data it can reach, and keep a human in the loop for risky actions. Guardrails that separate trusted instructions from unfamiliar content are the main defense.

New to all this? Start with what an AI agent really is, browse the full glossary, or explore the learning hub.

← Back to the glossary