Prompt Injection

Prompt injection is when an attacker injects instructions disguised as system instructions into the context via external content (web pages, emails, documents), hijacking the agent’s behavior.

More dangerous in agent systems than in chatbots: an agent can perform irreversible actions.

Context-level protection:

  • marking the source of external content;
  • strict message roles (separation of instructions and data);
  • input sanitization.

But this is only the first line of defense – protection at the execution level is also needed.

Related: Sandboxing, Deadly Triad, Agent Loyalty