TL;DR: Prompt injection is an emergent property of how large language models process context, not a conventional software flaw, and red teaming can surface failure modes without eliminating the risk, according to Noma Security. The governance task is to reduce blast radius, harden permissions, and accept that certainty is unavailable.
At a glance
What this is: Prompt injection is framed here as an inherent context-processing risk in generative AI, not a conventional bug, and red teaming is positioned as a way to expose failure modes rather than eliminate them.
Why it matters: IAM and security teams need to treat GenAI controls as governance and blast-radius problems, because unsafe model output can still drive downstream tool use, data exposure, and permission misuse.
Context
Prompt injection is the security failure mode that appears when a model treats attacker-supplied text inside its context window as instruction rather than data. The article argues that this is not a conventional software bug, because the surrounding application is assembling prompts from documents, history, and tool metadata before the model responds.
For identity and access teams, the important question is not whether prompt injection can be made impossible. It is whether GenAI deployments are being given permissions, tool reach, and data access that make a single malformed instruction able to move beyond text generation into operational action.
That shifts the problem from model quality to governance design. If the model can influence search, document updates, or external actions through downstream tools, then the control objective is to contain the blast radius of a bad instruction, not to assume the model itself will reliably reject it.
Key questions
Q: What breaks when prompt injection is not governed like an access problem?
A: The organisation may treat malicious text as a harmless message, even though it can steer an agent into exposing data or taking privileged actions. Prompt injection is dangerous because it turns untrusted content into a control plane for behaviour. Teams need policy and authorisation checks around outputs, not just message filtering.
Q: Why do GenAI systems need blast-radius controls if red teaming is already in place?
A: Because red teaming can reveal failure modes, but it cannot make probabilistic model behaviour deterministic. Blast-radius controls limit the damage when a model follows malicious instructions or misinterprets benign ones. Without tight permission scope, the test results may improve while the real-world exposure remains high.
Q: What are the signs that an AI deployment is too permissive for prompt injection risk?
A: Look for models that can read broad internal content, write to shared documents, or trigger workflows without a separate authority check. If a bad instruction in retrieved content can reach a sensitive tool path, the deployment is overexposed. The danger is not the prompt alone but the trust granted to its outputs.
Q: How should security teams run AI red teaming for GenAI systems?
A: Start with the system’s trust boundaries, then test prompts, retrieval sources, tool calls, and output handling together. Good AI red teaming does not stop at finding bad answers. It checks whether the system can be pushed into revealing data, bypassing policy, or taking unsafe actions through connected interfaces.
Technical breakdown
Why the context window creates prompt injection exposure
Prompt injection works because many AI applications mix system instructions, user prompts, retrieved content, tool descriptions, and prior conversation into one text stream. The model does not see a hard security boundary between trusted instruction and untrusted data. It predicts the next likely output from the full context, which means malicious instructions embedded in documents can be interpreted as relevant content rather than hostile input.
Practical implication: Treat every retrieval and tool input as potentially instruction-bearing, and design for containment rather than trust in prompt separation.
Why nondeterminism makes AI security probabilistic
Large language models are not deterministic. The same prompt can produce different outcomes as phrasing changes, randomness shifts, or model updates land. That makes prompt injection a risk that can be reduced but not eliminated, because any safety rule is being inferred by the model rather than enforced by a hard policy boundary. A guardrail can work in one run and fail in the next, even when the visible input looks similar.
Practical implication: Assume repeated testing will reveal risk patterns, not produce a guarantee, and place limits on what the model can do when interpretation varies.
How downstream tools turn model output into operational risk
Prompt injection becomes materially dangerous when another system acts on the model's output. The model may simply generate text, but the surrounding application can use that text to search data, update documents, or trigger workflows. In that case, the security boundary is no longer the model itself. It is the combination of model output, tool permissions, and the trust granted to the orchestrating application. This is where mis-scoped permissions and unsafe defaults become incident drivers.
Practical implication: Constrain tool permissions and separate text generation from execution authority so model output cannot directly drive sensitive actions.
Breaches seen in the wild
- DPD chatbot incident 2024: A customer got DPD's AI chatbot to swear and mock its owner; DPD blamed an error after a system update and disabled the AI element.
- 12,000 secrets in LLM training data: Truffle Security found 11,908 live API keys and passwords hard-coded in web pages captured by Common Crawl, a dataset used to train LLMs.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Prompt injection is a governance failure mode because the control boundary sits around the application, not inside the model. The article makes clear that the model is doing what it was trained to do when it follows language in context, even if that language is malicious. The failure is the assumption that guardrails inside the prompt can substitute for enforceable boundary control. Practitioners should treat prompt injection as a design and authority problem, not a model bug.
Blast-radius control is the decisive security variable in GenAI deployments. Red teaming can surface unsafe defaults, mis-scoped permissions, and fragile prompts, but it does not convert probabilistic behaviour into certainty. That means the governance question is how much damage a successful injection can cause when tool access, data access, or write privileges are too broad. The practical implication is that security posture depends more on containment than on perfect detection.
AI red teaming is valuable precisely because it proves uncertainty, not safety. The article correctly rejects the idea that testing can eliminate prompt injection, while also showing why testing still matters for responsible deployment. This positions red teaming as an input to governance decisions about acceptable exposure, not a final assurance mechanism. Organisations that expect definitive pass or fail answers from red teams are asking the wrong question.
Prompt injection exposes a context-trust debt across GenAI programmes. The surrounding application bundles documents, history, and instructions into one execution context, but governance often still assumes those inputs are separable. That assumption fails when attacker-controlled text can steer model output and then reach downstream tools. The implication is that AI governance has to account for context construction, permission scope, and execution authority as one control plane.
Responsible GenAI adoption will increasingly be measured by constrained authority, not model confidence. The article shows that even well-tested systems remain probabilistic, which means leaders need to judge deployments by how tightly they limit what a model can access and do. That shifts the discipline from asking whether the model is safe in the abstract to asking whether the surrounding access model makes unsafe outcomes containable.
From our research library:
- AI-related credential leaks surged 81.5% year-over-year in 2025, with the surrounding AI infrastructure leaking 5x faster than core LLM providers, according to the State of Secrets Sprawl 2026.
- Read next: Agentic AI Security Guide
What this signals
Context-trust debt: GenAI systems often treat retrieved content, prior dialogue, and tool metadata as one execution context, which makes trust placement the real control problem. Security leaders should expect the next wave of governance to focus on what the model is allowed to influence, not just what it is allowed to see.
Red teaming matters because it identifies the conditions under which a model will follow malicious instructions, but those findings only become useful when mapped to authority limits. The practical programme response is to align prompt design, retrieval scope, and tool permissions so that a single successful injection cannot become an operational event.
For practitioners
- Constrain downstream tool permissions Limit what the model can search, write, update, or trigger so a successful injection cannot directly reach sensitive data or business workflows.
- Separate retrieval from execution Treat retrieved documents and conversation history as untrusted input, then keep execution privileges outside the prompt construction layer.
- Red-team realistic prompt variants Test indirect instructions, multilingual rephrasing, and ambiguous business language to see where the system follows malicious guidance.
- Review unsafe defaults before scale-up Look for permissive tool settings, broad data access, and fragile system prompts before expanding the deployment to more users or workloads.
Key takeaways
- Prompt injection is best understood as a governance and authority problem, not a conventional software defect that can be patched away.
- Red teaming helps teams see where GenAI systems fail, but it cannot turn probabilistic behaviour into a guarantee of safety.
- The effective control is blast-radius reduction through narrow tool permissions, careful retrieval handling, and clear separation between model output and execution.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | Prompt injection becomes dangerous when model output drives unsafe tool actions. |
| ASI09 — Human-Agent Trust Exploitation | The article is fundamentally about language manipulation that exploits trust in AI output. | |
| Recommendation — Restrict tool use so model output cannot trigger sensitive actions without separate authorization. Assume model-generated text can be socially engineered and validate it before execution. | ||
| NIST AI RMF | GOVERN — AI Governance and Accountability | The article frames prompt injection as a governance issue for deployment decisions. |
| Recommendation — Assign clear accountability for GenAI risk acceptance, testing, and operational limits. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | Least-privilege authority is central to limiting the impact of prompt injection. |
| Recommendation — Apply least-privilege authorization to every AI-connected workflow and downstream tool. | ||
| MITRE ATT&CK | TA0006;TA0010 — Credential Access; Exfiltration | The article discusses malicious model steering that can lead to data access and leakage. |
| Recommendation — Map AI-assisted data exposure paths to credential access and exfiltration tactics. | ||
Key terms
- Prompt Injection (Agentic): An attack where malicious instructions are embedded in content that an AI agent reads, causing the agent to execute unintended actions using its own legitimate credentials. A primary vector for agent goal hijacking and identity abuse.
- Context Window: The context window is the text a model receives at one time, including prompts, retrieved documents, and conversation history. Security teams care about it because it becomes the practical boundary between trusted instructions and untrusted content, especially when the application assembles that text automatically.
- AI Red Teaming: AI red teaming is the practice of simulating hostile behaviour against models, applications, and agents to expose weaknesses before real attackers do. In AI programmes, it is most useful when results can be turned into controls, monitoring, and governance evidence rather than left as a one-time test report.
- Blast Radius: The potential scope of damage if a specific credential or identity is compromised. Identities with broad permissions have a larger blast radius and represent a higher priority for least-privilege enforcement and security controls.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on May 30, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org