← Back to Intelligence Library

Labs & Projects • Agentic AI Security Research

Indirect Prompt Injection: The Risk in What Agents Read

A study of malicious instructions in retrieved documents, tool responses, and agent memory, with a proposed synthetic-data defense lab.

Agentic AIIndirect Prompt InjectionRAGTrust Boundaries

Last Updated: September 28, 2026

Research scope

This article adapts the Indirect Prompt Injection research note. Its defense lab and evaluation metrics are proposals for controlled testing, rather than results from a completed project.

The attacker controls the information environment

An agent may encounter malicious instructions in a webpage, support ticket, email, repository comment, retrieved document, or tool response. The person requesting help can be acting legitimately while an unrelated source attempts to redirect the workflow.

The central failure is promoting lower-trust content into authority. Successfully retrieving a document establishes that the tool worked; it does not make instructions inside that document authorized.

Memory and handoffs extend the risk

The note also considers persistence and propagation. Poisoned context may enter a retrieval collection or persistent memory and affect a later task. A summary passed to another agent can carry the same manipulation even when that agent never reads the original material.

Preserving provenance matters at each boundary. A handoff should not silently turn an external claim into a trusted instruction.

Controls at the action boundary

Constrain tool availability to the task, check resource and destination permissions independently, and prevent sensitive information from automatically flowing into external messages. Trust labels help explain provenance, but enforcement must occur in the application around the model.

Logging should connect the original objective, retrieved resources, requested tool operations, approvals, and final actions. Behavioral changes are more informative than looking only for a particular suspicious phrase.

What the proposed lab would measure

The source proposes using synthetic documents and a harmless test secret to compare an unrestricted agent with a protected configuration. Measurements include attack success, unauthorized tool requests, authorization blocks, disclosure of synthetic data, and detection.

These measurements answer different questions. An agent requesting a forbidden operation is a model-level failure even if a policy check successfully blocks execution. Recording both outcomes makes the value and limitations of the control visible.

Practical takeaway

Treat external content as evidence to evaluate. It cannot expand the task or grant new permissions simply because an agent encountered it during useful work.

Source research

Adapted from Indirect Prompt Injection. Consult the original note for its references, evidence, and full analysis.

Millie's Perspective

A legitimate retrieval tool can return hostile content. Source provenance and task-scoped permissions should remain visible throughout retrieval, planning, and tool execution.

Key Takeaways

  • The attacker can influence an agent without using its chat interface.
  • Retrieved content and tool responses do not grant authority.
  • Memory and agent handoffs can carry poisoned context forward.
  • Measure attempted actions, blocked actions, and actual disclosure separately.

Project Repository

Interested in the complete project, lab documentation, or research notes? Explore the full repository on GitHub.

View on GitHub →

Related Reading