Because they blend untrusted content and trusted instructions inside the same execution path. If logs, files, binaries, or retrieved documents can influence the model before the harness checks them, attacker-authored content can steer tool use, which turns a content issue into an access-control issue.
How agentic workflows turn prompt injection into an execution problem
Agentic workflows widen the attack surface because the model is not just reading text, it is acting on it. Once retrieved content, logs, files, or web pages can influence planning, the model can convert attacker-authored instructions into tool calls, data access, or workflow changes before a human reviews the result. That makes the problem less like prompt quality and more like unsafe trust propagation across a control boundary.
The key shift is that the model often sits inside a trusted orchestration path. Content that would be harmless if treated as data becomes dangerous if the harness allows it to compete with system instructions, memory, or tool context. In practical terms, the same input channel can carry both information and command-like influence, so the workflow must assume that any untrusted content may try to steer the next action.
That is why indirect prompt injection is usually more serious in agentic systems than in chat-only systems. A chat model can produce a bad answer, but an agent can search, summarize, send, approve, delete, or purchase. When the model’s output is wired to real privileges, the blast radius includes whatever those tools can reach.
Where the risk comes from in retrieval, logs, and tool use
Indirect prompt injection works when attacker content enters the agent’s context through a path the system treats as safe, such as retrieval-augmented documents, tickets, emails, notes, binaries, or telemetry. If the agent consumes that content before policy checks, the attacker can hide instructions inside seemingly ordinary data and exploit the model’s tendency to follow the most salient local cue.
This is especially hazardous in pipelines that merge multiple sources into one prompt. If the workflow does not preserve source trust levels, the model may not know which text is authoritative, which text is adversarial, and which text is only evidence. The failure mode is not that the model “gets confused” in a generic sense, but that untrusted material is allowed to participate in authorization-sensitive decision-making.
Engineered agents also create a second risk: tool routing can amplify a small injection into a larger compromise. A malicious document can ask the agent to fetch more data, expose a secret, or call a connector that was never intended for that document. NHIMG’s Agentic AI Security Guide treats prompt injection as part of a broader agent threat model because the practical harm comes from the chain between input, planning, and action. The same pattern is reflected in the OWASP Agentic AI Top 10, where goal hijacking, tool misuse, and identity and privilege abuse sit close together.
Why the harness, privilege model, and containment strategy matter
The security boundary is not the model alone, it is the harness around the model. If the harness allows retrieved content to reach planning state before validation, or if tools are callable without per-action checks, indirect prompt injection can become an access-control failure instead of a content moderation failure. That is why agent design must separate observation, interpretation, and execution as much as possible.
Practical containment starts with least privilege, scoped credentials, and explicit approval gates for sensitive actions. An injected instruction should never be able to inherit broad workspace or human session authority just because it appeared in context. AI Agent Authorisation Guide maps that problem to task-scoped access and per-action policy, while Zero Trust for AI Agents frames the same need as verifying the principal, request, and privilege before every meaningful action.
Containment also matters for browser and desktop agents, because the model may inherit live sessions, cookies, and page state. Browser and Computer-Use Agent Security Guide is useful here because the workflow risk rises sharply when untrusted content can steer an agent that already has authenticated access to a user’s environment.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Agent prompts can steer privileged tool use and access decisions. |
| ASI02 — Tool Misuse | Indirect injections exploit agent tools to turn content into unsafe actions. | |
| ASI01 — Agent Goal Hijack | Injected instructions can redirect the agent away from its intended task. | |
| Recommendation — Enforce per-action authorization and limit agent privileges to the minimum needed. Restrict tool invocation to validated intents and approved workflows. Validate task intent before execution and reject conflicting instruction sources. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Limiting tool and session privilege reduces blast radius from injected prompts. |
| AC-3 — Access Enforcement | Execution gates are needed so untrusted content cannot directly trigger actions. | |
| AU-2 — Event Logging | Tracing agent inputs and tool calls is essential for detecting injected paths. | |
| Recommendation — Grant agents only the access required for the current task. Enforce authorization checks before every sensitive agent action. Log agent inputs, decisions, and tool actions for review and response. | ||
| NIST Zero Trust (SP 800-207) | Zero Trust Architecture | The topic depends on verifying each request and not trusting context by default. |
| Recommendation — Verify each request and assume untrusted content may be hostile. | ||
Practitioner Guidance
What to verify: Check whether untrusted sources are separated from executable instructions at the harness layer, not just in the prompt template. If retrieved text, log lines, or email content can influence tool selection before a trust decision is made, you have a control gap, not a model quirk.
Decision rule: If the agent can take an action that changes state, reaches external systems, or exposes data, treat any untrusted content path as potentially adversarial and require an explicit policy check before execution. If the action is read-only and sandboxed, the residual risk is lower, but the workflow still needs source labeling and containment.
What good looks like: The agent can read untrusted material without granting that material any authority. Tool calls are scoped, logged, and attributable, and high-impact actions require a separate confirmation path that the injected content cannot satisfy on its own.
Practitioner takeaway: Indirect prompt injection becomes dangerous when the workflow collapses “what the agent reads” and “what the agent is allowed to do” into the same trust channel. The main control objective is to break that linkage so untrusted content can inform analysis without being able to authorise action.
Related resources from NHI Mgmt Group
- Why does indirect prompt injection increase risk for AI assistants in enterprise inboxes?
- Why do indirect prompt injection attacks create more risk in RAG and agentic applications?
- Why do agentic browsers increase phishing and prompt injection risk?
- Why do AI agents increase non-human identity risk in existing IAM programmes?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org