Prompt injection
Level 3
Hiding instructions inside content so a model obeys the attacker instead of you.
Term 4 of 6 in Safety and privacy
In plain language
If a model reads a web page or email, any text in it can be read as an instruction. Hidden text can tell an agent to leak data or take an unwanted action.
Think of it like this
Slipping a forged note into a courier's pouch and having it followed without question.
Why it matters
Prompt injection is the main security risk of AI agents, and the reason autonomous tools need strict limits on what they can reach.
Next terms in Safety and privacy
- BiasA model systematically favouring or disadvantaging certain people or outcomes.
- Personally identifiable informationAny information that can be traced back to a specific person.
- Data retentionHow long a service keeps what you send it, and what it does with it.
- AlignmentHow well an AI system's behaviour matches what people actually intend.