Labs & Projects • Agentic AI Security Research
Prompt Injection: When Input Becomes an Instruction
Last Updated: September 28, 2026
Research scope
This article adapts the Prompt Injection note in the Agentic AI Security Lab. The note is marked RESEARCHING and proposes a defensive experiment; it does not document a completed implementation or measured lab results.
The trust boundary
Prompt injection occurs when a language model treats untrusted input as an instruction that changes its intended behavior. Direct injection enters through an interaction with the model; indirect injection arrives in material the system reads while carrying out a legitimate request.
For a tool-enabled agent, the consequences can extend beyond an incorrect answer. A changed objective may lead to inappropriate file access, an unauthorized message, or a modification to a connected system. The available permissions determine how much damage that decision can cause.
Why a better prompt is not enough
The research centers on the boundary between proposing an action and authorizing it. A model can generate plausible tool arguments without having permission to use that tool on a particular resource. Independent policy checks must make that decision.
Defensive design priorities
Give each workflow only the permissions and tools it needs. Keep retrieved content separate from trusted instructions, validate proposed arguments and destinations, and require review before sensitive actions. Log requested actions alongside authorization decisions and actual outcomes so a later investigation can reconstruct the sequence.
A proposed comparison lab
The note proposes a small agent with webpage retrieval, an internal test note, and a messaging tool. In an isolated environment, synthetic hostile content would attempt to redirect the agent away from its assigned task.
Compare a baseline configuration with least privilege, independent tool authorization, and approval requirements. Use harmless test data, record attempted and completed actions separately, and repeat the same scenario after each control change. These are proposed evaluation steps, not reported successes.
Practical takeaway
Design controls for the possibility that the model misunderstands a source. A robust authorization boundary should still prevent an out-of-scope action after that misunderstanding.
Source research
Adapted from Prompt Injection. Consult the original note for its references, evidence, and full analysis.
Millie's Perspective
The useful design question is what an agent can do after it misinterprets input. Restricting permissions and checking actions independently makes that question testable.
Key Takeaways
- Untrusted content can redirect an agent toward an unauthorized objective.
- Tool access increases the consequences of manipulated reasoning.
- Authorization must be enforced independently of the model.
- The proposed lab compares controls; completed test results are not claimed.
Project Repository
Interested in the complete project, lab documentation, or research notes? Explore the full repository on GitHub.
View on GitHub →