TL;DR: Indirect prompt injection succeeds when malicious instructions are embedded in trusted data and LLMs can act on them across sensitive workflows, according to Pillar Security’s analysis. The real risk is not the payload alone but the combination of private data access, untrusted inputs, and external communication that turns prompt attacks into operational exploits.
At a glance
What this is: This analysis explains why indirect prompt injection succeeds when malicious instructions are embedded in trusted content that an LLM can process and act on inside business workflows.
Why it matters: IAM and security teams need to treat LLMs as workflow participants whose access to data, tools, and outbound channels can turn untrusted content into an execution path.
Context
Indirect prompt injection is a workflow security problem, not just a prompt-writing problem. It appears when a model processes untrusted content and treats embedded instructions as authoritative enough to influence output or actions.
In identity terms, the risk rises when LLMs can reach private data, handle external inputs, and send information out of the environment. That combination turns ordinary business content such as emails, tickets, documents, or code comments into a potential control plane for misuse.
Key questions
Q: What breaks when an LLM application treats untrusted content as instruction?
A: Prompt injection works because the application collapses the line between data and control. Once retrieved text, memory, or user content can influence privileged prompts, the model may follow attacker-supplied instructions, expose sensitive context, or trigger tool calls that were never intended by the operator.
Q: Why do private-data access and outbound tools make prompt injection worse?
A: Because prompt injection becomes operational when the model can read something valuable and send it somewhere useful. Private-data access creates the target, untrusted inputs create the vector, and outbound tools create the exfiltration path. Remove any one of those conditions and the attack loses force.
Q: How can security teams tell whether an LLM workflow is high risk for indirect prompt injection?
A: Look for workflows that combine external inputs, privileged context, and any ability to communicate outside the environment. If the model can consume untrusted content while also seeing sensitive data or using tools, the workflow deserves higher control, tighter scoping, and explicit review boundaries.
Q: Should organisations compare prompt filtering with workflow isolation for LLM security?
A: Prompt filtering helps, but it does not replace workflow isolation. Filtering tries to recognise malicious text after it arrives, while isolation reduces the chance that the model can turn that text into a harmful action. The safer design decision is to reduce trust exposure first, then add filtering as a secondary layer.
Technical breakdown
Why indirect prompt injection works in LLM workflows
Indirect prompt injection succeeds because the model cannot reliably distinguish user intent from instructions hidden inside the data it is asked to process. The attack does not need to break the model’s weights or bypass authentication. It only needs the model to treat embedded content as relevant context while the surrounding workflow gives it access to private data or tools. That makes the exploit operational, because the harm comes from the model following instructions inside normal business data, not from a visible malicious command surface. The CFS model in the article captures this by showing how context, format, and salience shape whether the payload is ignored or executed.
Practical implication: isolate untrusted inputs from instructions before they reach LLMs that can access sensitive systems.
Why context, format, and salience change exploitability
The article’s CFS model explains why some injections succeed and most fail. Context means the payload matches the task the model is performing. Format means it looks native to the medium, such as an email, code comment, or HTML block. Salience means it is positioned and phrased to attract attention, often with imperative language and explicit instructions. These three factors matter because LLMs are highly sensitive to what appears relevant, authoritative, and well formed. An injection that fits the workflow is far more likely to be followed than one that feels random or out of place.
Practical implication: review high-risk workflows for content formats where malicious instructions can blend into expected structure.
Why the lethal trifecta makes prompt injection operational
The article leans on Simon Willison’s lethal trifecta to show when indirect prompt injection becomes genuinely dangerous: access to private data, exposure to untrusted content, and external communication ability. Each capability alone is useful; together they create a data exfiltration path. The model can read sensitive information, consume attacker-controlled input, and then leak data through a response, email, ticket, or other outbound channel. That is why the issue is not the prompt alone. The exploit emerges from the surrounding identity and workflow design that lets the model bridge trust boundaries.
Practical implication: map every LLM workflow that combines private data, untrusted input, and outbound communication as a high-risk identity path.
Threat narrative
Attacker objective: The attacker wants the LLM to exfiltrate sensitive data or perform an unauthorized action through a trusted business workflow.
- Entry occurs when attacker-controlled content is placed into a trusted channel such as email, documentation, support tickets, or source files.
- Credential or data access follows when the LLM reads sensitive context or has access to connected tools and private records during its task.
- Escalation happens when the hidden instruction redirects the model to reveal data, query restricted sources, or misuse its outbound communication path.
- Impact is data leakage or unauthorized action performed through a workflow that appeared legitimate to the human reviewer.
Breaches seen in the wild
- Gemini CLI prompt injection flaw 2025: Tracebit showed a poisoned README could make Gemini CLI run hidden commands and exfiltrate developer secrets; Google fixed it in 0.1.14.
- EchoLeak (Microsoft 365 Copilot) 2025: A crafted email could make Microsoft 365 Copilot leak data from its context with no click, a zero-click prompt injection fixed as CVE-2025-32711.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
Indirect prompt injection is becoming a workflow integrity problem before it becomes a model-safety problem. The article shows that the exploit succeeds by abusing the trust boundary between content and instruction inside operational workflows. That makes the governing question less about whether the model is smart enough and more about whether the surrounding process lets untrusted data become action-bearing input. Practitioners should treat content-to-action pathways as identity-sensitive control points.
The lethal trifecta is the real governance boundary for LLM deployments. Access to private data, exposure to untrusted content, and outbound communication ability form a compound risk condition that changes the security model. Remove any one of those conditions and the exploitability drops sharply. That is why governance has to evaluate the full workflow, not just the prompt surface, when deciding whether an LLM can participate in a business process.
Context, format, and salience describe a repeatable attack design pattern, not an edge case. The article’s most useful contribution is showing why injections work when they are embedded in the medium the model already expects, placed where attention is highest, and written to look authoritative. That means defenders need to understand how attackers optimize for model attention, not just whether a payload is obviously malicious. Security teams should assume that well-formed hostile content will keep getting better.
Prompt injection governance will converge with identity and access governance. Once an LLM can read, decide, and act across systems, the trust decision is no longer only about content moderation. It becomes about which data, tools, and outbound channels the workflow is allowed to touch. The organizations that will handle this best are the ones that already think in terms of scoped access, explicit boundaries, and revocation paths rather than implicit trust.
High-risk AI workflows need a named concept for the combined exposure surface. Lethal trifecta exposure: the condition in which private data access, untrusted inputs, and external communication coexist in one LLM workflow. That combination turns prompt injection from nuisance into operational exploit. Practitioners should use that concept to prioritize which AI-enabled processes deserve immediate governance review.
From our research library:
- 96% of organisations store secrets outside of secrets managers in vulnerable locations including code, config files, and CI/CD tools, according to the Ultimate Guide to NHIs.
- 88% of organisations have embedded AI agents in their workflows, according to KPMG's 2026 report.
- Read next: Agentic AI Identity Guide
What this signals
Lethal trifecta exposure: the combination of private data access, untrusted content, and outbound communication is the condition that should drive your AI governance triage. When those three coexist, the control problem moves from prompt hygiene to workflow containment and privilege design.
Pillar Security’s analysis also points to a larger programme shift: LLM risk reviews need to sit alongside identity, data access, and tool authorization reviews. A model that can read sensitive material and send it elsewhere is no longer just a generative system; it is an access path that needs scoped trust, containment, and revocation.
For practitioners
- Define lethal-trifecta workflows Inventory every LLM workflow that combines private data access, untrusted inputs, and outbound communication. Those are the places where indirect prompt injection can become operational rather than theoretical.
- Separate instructions from content Design processing steps so the model can read untrusted material without treating it as instruction-bearing text. Use structural separation, content normalization, or human review where content and control signals currently mix.
- Limit tool and data scope Reduce what the model can query, send, or reveal in any workflow that consumes external content. The smaller the allowed action set, the less damage a hidden instruction can cause.
- Inspect high-salience input channels Prioritise email, ticketing, documentation, and code-comment paths where attacker text can sit at the beginning or end of the payload and still appear native to the workflow.
- Treat outbound communication as a control Require explicit review or containment for any LLM task that can send information outside the environment, especially when the task also has access to private records or inboxes.
Key takeaways
- Indirect prompt injection succeeds because trusted workflows let attacker-controlled content behave like instructions inside the model’s operating context.
- The article’s real warning is that private data access, untrusted inputs, and outbound communication together create a practical exfiltration path.
- Security teams should govern LLM workflows by reducing trust exposure, scoping tool access, and isolating content from control signals.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI09 — Human-Agent Trust Exploitation | The article is about untrusted content steering LLM behaviour through trusted workflows. |
| ASI02 — Tool Misuse | The attack becomes harmful when the model uses connected tools or outbound channels as part of the exploit. | |
| Recommendation — Map prompt injection risks to ASI09 and separate trusted instructions from attacker-controlled content. Constrain tool access so injected instructions cannot trigger sensitive actions or exfiltration. | ||
| NIST AI RMF | GOVERN — AI Governance and Accountability | The article is fundamentally about governing AI workflow trust, scope, and accountability. |
| Recommendation — Define AI governance boundaries for data access, input trust, and action authority before deployment. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | The key issue is overbroad LLM access to data and communication channels. |
| Recommendation — Apply PR.AA-05 to narrow what LLM-enabled workflows can read, query, and send. | ||
Key terms
- Indirect Prompt Injection: Indirect prompt injection is an attack where malicious instructions are hidden inside content that an AI system reads later. The model may treat that content as context rather than as hostile input, which can influence tool use, data access, or workflow actions if controls are weak.
- Lethal Trifecta: A risky AI agent condition where one system can read private data, consume untrusted content, and communicate externally. When those three capabilities overlap, the agent can be tricked into disclosing sensitive information through legitimate tools without a conventional exploit.
- Instruction Salience: The degree to which a hidden instruction stands out to an LLM and is likely to be prioritised during inference. Placement, wording, and apparent authority all affect salience, which means attackers often tune both content and position to make the model more likely to comply.
- Context-Format-Salience Model: A way to explain why some indirect prompt injections work and others fail. Context is whether the payload matches the task, format is whether it blends into the medium, and salience is whether it captures the model’s attention strongly enough to influence action.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
Published by the NHIMG editorial team on June 11, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org