TL;DR: Prompt injection is the top OWASP LLM vulnerability because attackers can override model behavior with plain language, and indirect injections in documents or retrieved data can redirect chatbots and agents without code exploits, according to WitnessAI. The real failure is assuming natural-language systems can be governed like structured applications when their input and instruction boundaries are not technically separable.
At a glance
What this is: This analysis explains why prompt injection lets attackers steer LLMs through natural language, with indirect injections posing the bigger enterprise risk when retrieved content or documents become attack carriers.
Why it matters: It matters because IAM, NHI, and AI governance teams cannot rely on classic input boundary controls alone when model instructions, user prompts, and downstream actions share the same execution context.
Context
Prompt injection is a control-boundary problem in AI systems: the model cannot reliably separate trusted instructions from untrusted text when both are processed as the same stream. That matters for identity security because the same weakness affects customer-facing chatbots, internal copilots, and autonomous agents that can act on what they read.
The governance gap is not limited to output quality. Once an AI system can consume documents, retrieve data, call APIs, or trigger workflow actions, a malicious instruction can turn ordinary content into a control input. In that sense, prompt injection is a shared-responsibility failure across model providers, application teams, and the enterprise programme that authorises the workflow.
Key questions
Q: What breaks when prompt injection protections are missing in AI-enabled security workflows?
A: Without prompt injection protections, attackers can manipulate the model into ignoring safeguards, revealing sensitive data, or producing unsafe actions and outputs. That can undermine trust in AI-assisted workflows, especially where the model has access to internal systems or security data. Teams should assume hostile inputs, constrain tool use, and filter both prompts and outputs.
Q: Why does prompt injection create greater risk in enterprise GenAI than in public-facing chatbots?
A: Enterprise GenAI usually has access to proprietary data, internal workflows, and customer information, so a successful prompt injection can do more than produce a bad answer. It can expose sensitive data, disrupt business processes, and trigger compliance failures. Public-facing systems are still vulnerable, but the enterprise blast radius is typically larger because the model sits closer to core operations.
Q: What are the signs that an AI agent may be vulnerable to prompt injection?
A: Look for mismatches between the prompt a system received and the actions it attempted, especially unexpected data retrieval, unusual API calls, or tool use that does not match the user's request. Those are strong indicators that input steering is affecting execution.
A: Teams should place enforcement at the action layer, not only inside the model. Guardrails can shape text and refuse many bad prompts, but they remain probabilistic and advisory. Runtime governance should check authorization, policy, and audit at the exact moment a tool call is made, so a manipulated agent can be stopped before money moves, data is deleted, or records are changed.
Technical breakdown
Why prompt injection works without code exploits
Prompt injection manipulates an LLM by using natural language to override the system prompt, not by exploiting a memory corruption flaw or malformed packet. The model receives system instructions, user input, and retrieved content as one token stream, so it has no native boundary equivalent to parameterised queries in SQL. That is why direct attacks can be typed into chat, while indirect attacks can hide inside documents, emails, or knowledge base entries. The model is doing what it was built to do: continue the conversation. The problem is that conversation itself becomes the attack surface.
Practical implication: treat every text-bearing input path into an LLM as a potential control surface, not just a content channel.
Why indirect prompt injection is the enterprise problem
Indirect prompt injection is more dangerous than a simple chat prank because the malicious instruction travels inside content the model is expected to trust and summarise. In retrieval-augmented generation and agentic workflows, that content can be a document, ticket, email, or knowledge base article that the model later consumes while performing a task. Once the model accepts the hidden instruction, it may reveal sensitive data, change tone, or issue tool calls as part of its normal operating flow. The risk grows when the model can chain actions across multiple systems, because one poisoned source can affect an entire workflow.
Practical implication: inspect retrieved and embedded content before it reaches the model, especially where the model can read and act in the same session.
Why runtime defence has to sit outside the model
Model-level safety training helps, but it does not create a reliable security boundary. The enterprise still decides what data the model can see, what actions it can trigger, and which connected systems it may call. That is the shared-responsibility gap: provider guardrails reduce risk in the model, while the enterprise must enforce policy at the point where prompts, responses, and actions intersect. Runtime inspection is therefore architectural, not cosmetic. It has to evaluate both inbound prompts and outbound responses before anything is exposed to users or downstream systems.
Practical implication: place policy enforcement in the runtime path so the control can block harmful prompts and prevent unsafe actions before execution.
Threat narrative
Attacker objective: The attacker wants the model or agent to ignore its intended boundaries and perform disclosure or actions that benefit the attacker.
- Entry occurs when an attacker submits malicious instructions directly into a chat interface or hides them inside a document, email, or retrieved record that the model will later process.
- Credential or data exposure follows when the model treats the injected text as authoritative and reveals system prompts, confidential content, or connected-system data it should not disclose.
- Escalation occurs when an agent accepts the injected instruction as part of its workflow and invokes APIs, queries databases, or triggers actions outside the user’s intent.
- Impact lands as exfiltration, unauthorised transactions, brand damage, or downstream operational harm caused by a model or agent that followed attacker-authored instructions.
Breaches seen in the wild
- Gemini CLI prompt injection flaw 2025: Tracebit showed a poisoned README could make Gemini CLI run hidden commands and exfiltrate developer secrets; Google fixed it in 0.1.14.
- Gemini AI Breach — Google Calendar Prompt Injection: Gemini AI assistant prompt injection attack leaks sensitive Google Calendar data.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Prompt injection is an instruction-boundary failure, not a content-filtering problem. The vulnerability exists because LLMs do not separate executable instructions from ordinary text in the way classic software separates code from data. That means security teams are not defending against a malformed payload but against text that the model may treat as operational intent. The practical conclusion is that governance has to focus on where instructions are allowed to enter and what they are allowed to trigger.
Indirect prompt injection turns trusted business content into a control plane. Once a model can read documents, tickets, and retrieved records, attackers no longer need a visible chat interface to influence behaviour. This is why the risk compounds in RAG and agentic workflows, where the model can both consume external content and take actions from it. The governance implication is that content provenance and action authority must be reviewed together, not as separate problems.
Shared responsibility breaks down when enterprises assume model alignment is a complete control. Provider safety layers reduce some risk, but they do not replace enterprise enforcement over data exposure, tool access, and action boundaries. That assumption fails even more sharply in autonomous workflows, where a malicious instruction can become a multi-step action chain. Practitioners need to recognise that model safety is only one layer of a broader runtime control stack.
Runtime inspection is the named concept that closes the AI interaction gap. The article’s central lesson is that prompt and response filtering must happen in motion, before the model can transform text into disclosure or action. That is the only place where policy can still distinguish legitimate work from adversarial steering. Security programmes should treat runtime defence as the enforcement point for AI interaction policy, not as an optional add-on.
Agentic AI extends prompt injection from bad output to bad execution. A chatbot that answers incorrectly is a governance issue; an agent that executes an injected instruction is an identity and authorisation issue. The moment a model can call APIs, modify records, or transmit data, the trust problem moves from language quality to delegated action. Practitioners should govern that delegation with the same seriousness they apply to privileged access in other systems.
From our research library:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
- Read next: AI Agent Observability, Audit and Incident Response Guide
What this signals
Prompt injection changes AI governance from content moderation to runtime policy enforcement. Teams that rely on prompt wording alone will miss the more important control question: what can the model see, what can it decide, and what can it execute? That means the programme needs a control plane for prompts, responses, and tool calls, not just model selection.
Indirect injection is the risk that most programmes underweight. A poisoned document, ticket, or retrieved record can influence an AI workflow even when no attacker touches the chat interface. That makes content provenance, retrieval boundaries, and action scoping part of the same governance decision.
Runtime defence has to be the enforcement layer, not the final detection layer. If the model can read and act in the same transaction, security controls must inspect prompts before the model sees them and responses before they trigger downstream behaviour. In practice, that shifts AI risk management from post hoc review to pre-execution control.
For practitioners
- Map every AI input path to a control owner Inventory chat interfaces, retrieval sources, document stores, and agent tool paths so each prompt source has an accountable owner and a defined policy boundary.
- Inspect retrieved content before model consumption Screen documents, emails, tickets, and knowledge base entries for hidden instructions or adversarial text before they are passed into summarisation or agent workflows.
- Constrain agent tool authority Limit which APIs, databases, and external endpoints an agent can invoke so injected instructions cannot expand its authority beyond the approved task.
- Deploy bidirectional runtime filtering Enforce pre-model prompt inspection and post-model response inspection so both inbound instruction abuse and outbound disclosure are blocked before execution.
- Red team indirect injection paths Test poisoned documents, multi-turn manipulation, encoded text, and retrieved-data attacks across every AI workflow that can read and act on enterprise content.
Key takeaways
- Prompt injection succeeds because LLMs do not naturally separate instructions from untrusted text, so the control problem is architectural rather than purely content-based.
- Indirect injections are the most consequential enterprise risk because trusted documents and retrieved data can become attack carriers inside AI workflows.
- The effective response is runtime enforcement that constrains prompts, responses, and tool use before an AI system can disclose data or trigger actions.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Prompt injection becomes a delegated-action problem when agents can misuse authority given by the enterprise. |
| ASI02 — Tool Misuse | Injected prompts steer agents into calling tools and endpoints they should not use. | |
| Recommendation — Restrict agent authority so injected instructions cannot expand identity or privilege beyond the approved task. Constrain tool access and validate every call path against task scope before execution. | ||
| OWASP Non-Human Identity Top 10 | NHI-10 — Human Use of NHI | Prompt injection often works by using people to relay instructions into non-human workflows and trust paths. |
| Recommendation — Separate human-authored text from machine-authorised actions so people cannot unintentionally confer authority. | ||
| NIST AI RMF | MANAGE — Manage AI Risks | Runtime defense and shared responsibility are AI risk management issues, not only model-safety issues. |
| Recommendation — Implement runtime policy enforcement for prompts, outputs, and downstream actions as part of AI risk controls. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | The article centres on limiting what an AI system can access and do after instruction abuse. |
| Recommendation — Apply least-privilege permissions to AI-connected systems so injected instructions cannot widen access. | ||
Key terms
- Prompt Injection (Agentic): An attack where malicious instructions are embedded in content that an AI agent reads, causing the agent to execute unintended actions using its own legitimate credentials. A primary vector for agent goal hijacking and identity abuse.
- Indirect Prompt Injection: Indirect prompt injection is an attack where malicious instructions are hidden inside content that an AI system reads later. The model may treat that content as context rather than as hostile input, which can influence tool use, data access, or workflow actions if controls are weak.
- Runtime Defense: Runtime defense is the set of controls that inspect, constrain, and stop unsafe AI behavior while the system is operating. For hospitality deployments, it covers both incoming prompts and outgoing responses, plus the tool calls that agents may trigger after a model decides to act.
- Shared Responsibility Model: A shared responsibility model divides security duties between the cloud provider and the customer. For NHI governance, the provider supplies the platform controls, but the organisation still owns configuration, privilege review, secret handling, monitoring, and lifecycle management of its identities.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity security are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org