By NHI Mgmt Group Editorial TeamBased on WitnessAI: “What is Indirect Prompt Injection and How Does It Work?” (March 6, 2026)

TL;DR: Indirect prompt injection lets attackers hide malicious instructions inside emails, documents, web pages, and knowledge bases that AI systems already trust, and agentic AI can turn that into unauthorized actions with system credentials, according to WitnessAI. Existing guardrails, pattern matching, and application controls provide only partial coverage because the model cannot reliably separate trusted instructions from untrusted content.


At a glance

What this is: Indirect prompt injection is an attack path that turns trusted content sources into AI control channels, with WitnessAI arguing that the risk escalates sharply when agentic systems can act on poisoned instructions.

Why it matters: IAM, NHI and AI security teams need to treat content ingestion, tool access and delegated execution as one governance problem because the attack crosses boundaries that point tools do not reliably see.


Context

Indirect prompt injection is a content-mediated attack against AI systems. Malicious instructions are hidden inside data the model will later consume, such as email, documents, web pages and knowledge bases, so the attack rides through normal retrieval and context-building paths rather than a visible chat prompt.

For identity and access governance, the problem is not only bad output. When an AI system can retrieve content, interpret it and then act with delegated permissions, the attack surface shifts from information integrity to execution authority, which is the point where AI security and NHI governance start to overlap.

That makes this a control-boundary problem as much as a model-safety problem. If an organisation treats prompts, retrieved content, agent tools and action approval as separate risks, it will miss the trust chain that indirect prompt injection exploits.


Key questions

Q: How should security teams reduce indirect prompt injection risk in AI systems?

A: Security teams should limit what AI systems can read, separate untrusted content from privileged actions, and apply least privilege to every connected agent. The strongest posture combines content filtering, allowlisted sources, short-lived sessions, and explicit approval for sensitive actions. If any one of those layers is missing, the attack path remains open.

Q: Why do native guardrails fail against prompt injection in AI agents?

A: Native guardrails often classify text rather than control execution, so they can miss attacks that manipulate the agent’s next action instead of its visible output. Prompt injection, encoded instructions, and multi-turn coercion exploit that gap. The practical answer is to enforce policy deterministically at runtime so the agent cannot carry out an unsafe tool call even when the text looks benign.

Q: What breaks when an AI agent acts on poisoned content?

A: The failure is not limited to a bad answer. Once the agent can call tools or use delegated credentials, poisoned context can become unauthorized execution, data exposure or policy override. The control failure is at the point where untrusted content is allowed to influence an identity-bound action.

Q: Should organisations keep human approval gates for high-risk AI actions?

A: Yes, when the action is irreversible, externally visible, or capable of changing production state. Human approval should be reserved for the highest-impact decisions, while lower-risk actions can be governed by pre-approved policy. That balance preserves speed without turning automation into uncontrolled execution.


Technical breakdown

How indirect prompt injection rides through retrieval-augmented generation

Indirect prompt injection works because large language models read instructions and data through the same text-processing path. In retrieval-augmented generation, the system pulls in external content to answer a request, but the model cannot reliably tell whether a sentence came from a trusted system prompt or from poisoned source material. Attackers exploit that ambiguity by planting hidden instructions in content the system already expects to trust. Once the malicious text enters the context window, the model may follow it as if it were legitimate policy, even though no user typed it into the interface.

Practical implication: inspect retrieved content before it reaches the model, not just user-entered prompts.

Why pattern matching and guardrails miss layered injection

The article shows why simple detection fails. Attackers can use delimiter spoofing, role hijacking, language switching, semantic rephrasing, invisible text and long-context placement to make the same malicious intent look different each time. Keyword filters only catch the obvious cases, while model guardrails often fail when the payload is indirect, encoded or buried deep inside a long context. The weakness is structural: the attack is semantic, but the defence is often lexical. That mismatch lets poisoned instructions survive multiple layers of inspection.

Practical implication: use intent-based classification and multi-stage inspection instead of relying on signatures alone.

Why agentic AI turns content poisoning into identity abuse

Agentic systems raise the stakes because the model is no longer only producing text. It can call tools, chain actions and operate with delegated credentials, so a successful injection can become an authorised-looking execution path. The article highlights persistent memory, tool chaining and multi-agent workflows as amplifiers that extend a single poisoned instruction into broader compromise. In identity terms, the issue is not just what the model says, but what it is authorised to do after it has been steered by untrusted content.

Practical implication: bind tool permissions, approval gates and high-risk action controls to the agent session, not just the underlying model.


Threat narrative

Attacker objective: The attacker wants to steer a trusted AI system into revealing information or carrying out actions that appear authorised because they execute through the system's own context and permissions.

  1. Entry occurs when an attacker plants malicious instructions inside an external data source that the AI will later retrieve, such as an email, document, web page or knowledge base entry.
  2. Credentialed propagation follows when the system ingests the poisoned content during normal retrieval-augmented generation and folds it into the model context.
  3. Execution happens when the model treats the embedded instructions as legitimate and follows them inside the same session or across chained tools.
  4. Impact occurs when the compromised model or agent exfiltrates data, manipulates tool actions, overrides policy or triggers unauthorized downstream actions with delegated credentials.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Indirect prompt injection is a content trust failure, not just a prompt-safety issue. The attack works because organisations still assume retrieved content is materially different from instructions once it enters the model context. That assumption breaks the moment external text can alter behaviour without a visible prompt change. The practitioner implication is that ingestion, context handling and execution approval now belong in the same governance chain.

Agentic AI turns indirect prompt injection into an identity problem. Once a model can act with delegated permissions, the question shifts from bad output to unauthorized execution under system authority. The model is not the identity, but it is now a decision path that can exercise identity-bound access. Practitioners should treat agent sessions, tool calls and delegated actions as governable identity events, not just model events.

Bidirectional inspection is the new runtime control plane for AI security. The article makes clear that input scanning alone is insufficient because poisoned instructions can surface in retrieved content, while output scanning is needed to catch leakage, instructions and policy-breaking responses before they are consumed downstream. The governance lesson is that AI controls must watch both what enters the model and what exits it, especially where agentic workflows exist.

Intent-based classification is the named concept that closes the gap left by pattern matching. The article shows why lexical filters fail against rephrased, encoded or multilingual payloads. Intent detection is the practical shift from asking what the text looks like to what it is trying to make the system do. Practitioners should use it as the decision layer above signatures, not as a substitute for them.

Continuous control is now the right operating model for AI identity risk. No single control layer can solve indirect prompt injection because the attack can enter through content, context, tools or outputs. That means governance must be continuous across the full session lifecycle, from retrieval through action approval and auditability. For practitioners, the measurable target is controlled execution under untrusted content, not perfect prompt hygiene.

What this signals

AI identity governance now has to cover the full data-to-action path. The practical risk is not isolated prompt abuse, but the way retrieved content, tool calls and delegated permissions combine inside one execution flow. Security teams should assume that any content source the model can read is also part of the control surface.

Direct prompt controls are no longer enough for enterprise AI. The article shows that indirect injection can arrive through ordinary business content, so the governance model has to move upstream to retrieval, context handling and output review. That is especially important where agents can chain tools and carry memory across sessions.


For practitioners

  • Implement bidirectional prompt and response scanning Inspect both incoming context and model outputs before they reach users, downstream tools or approval workflows. Use separate checks for user input, retrieved content, tool output and final model response so poisoned instructions and leaked content are intercepted at each boundary.
  • Adopt intent-based content classification Classify whether content is trying to instruct, extract, override or exfiltrate rather than relying on keywords alone. This is the control layer that catches semantic rephrasing, multilingual payloads and encoded instructions that signature filters miss.
  • Tokenize sensitive data inline Replace raw sensitive values with reversible tokens before they enter model context so an injected instruction cannot easily exfiltrate or misuse the original data. Keep tokenization active across prompts, retrieved content and responses to preserve workflow utility while reducing exposure.
  • Bind high-risk actions to explicit approval Require human review for system changes, financial actions and protected-data access when an agent is operating on untrusted content. Make approval a property of the action path, not a vague policy attached only to the application.
  • Discover and govern shadow agents and MCP servers Inventory agent sessions, connected tools and MCP servers so unmanaged automation does not inherit broad access by default. Attribute each agent action to a human owner and verify that exposed tools are within the organisation's approved scope.

Key takeaways

  • Indirect prompt injection turns ordinary enterprise content into an execution path when AI systems cannot distinguish trusted instructions from untrusted text.
  • The attack becomes materially more dangerous when agentic systems can act with delegated credentials, because compromised context can produce unauthorized real-world actions.
  • The most defensible control model is continuous and layered, with input inspection, intent-aware classification, inline tokenization, action gating and audit trails.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbusePoisoned content can steer an agent into abusing delegated identity and access.
ASI02 — Tool MisuseThe article centres on malicious instructions driving harmful tool calls and chained actions.
ASI07 — Insecure Inter-Agent CommunicationMulti-agent and tool-chain propagation extend the injection path across system boundaries.
Recommendation — Apply identity-bound action controls to stop injected content from driving unauthorized tool use. Constrain agent tool execution with approval gates and per-action authorization. Validate inter-agent messages and tool outputs before allowing downstream execution.
OWASP Non-Human Identity Top 10NHI-04 — Insecure AuthenticationDelegated system credentials let compromised AI actions appear authorized.
NHI-10 — Human Use of NHIHuman operators and shadow agents both appear in the identity chain behind model actions.
Recommendation — Bind delegated AI access to strong authentication and explicit action scoping. Attribute every agent action to a human owner and enforce accountable use of machine identities.
NIST SP 800-53 Rev 5IA-5 — Authenticator ManagementCredential scope and lifecycle are central when agents can act with delegated permissions.
Recommendation — Tighten authenticator lifecycle controls for AI agents and revoke overly broad access quickly.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe article focuses on limiting what agentic systems may do with inherited access.
Recommendation — Review AI entitlements so agent permissions stay limited to approved business actions.
MITRE ATT&CKTA0006;TA0008 — Credential Access; Lateral MovementSuccessful injection can lead to credential abuse and chained movement through connected tools.
Recommendation — Map prompt-injection scenarios to credential access and lateral movement detections in monitoring workflows.

Key terms

  • Indirect Prompt Injection: Indirect prompt injection is an attack where malicious instructions are hidden inside content that an AI system reads later. The model may treat that content as context rather than as hostile input, which can influence tool use, data access, or workflow actions if controls are weak.
  • Intent-based classification: Intent-based classification evaluates what a user or system is trying to do, not just what text or file is present. In AI governance, it distinguishes routine work from risky interaction by reading context, purpose, and sensitivity. That matters when regulated data is handled conversationally rather than through formal file transfer.
  • Bidirectional Inspection: Bidirectional inspection means examining both prompts sent to an AI system and the responses it produces. This is essential when outputs can trigger follow-on actions, leak sensitive data, or carry instructions that affect downstream systems, users, or automated workflows.
  • Delegated Identity: Delegated identity is when one actor acts on behalf of another with explicit permission and bounded authority. In AI-assisted commerce, it requires clear consent, limited scope, and traceable records so the retailer can distinguish authorised delegation from unauthorised automation.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org