Cross-agent prompt injection occurs when one agent manipulates another through shared context, tool output or orchestration layers. It is especially relevant in multi-agent systems because trust is transferred between components, allowing malicious instructions to spread across otherwise separate workflows.
What Cross-Agent Prompt Injection Is
Cross-agent prompt injection is not just a single-agent jailbreak. The defining feature is that malicious instructions move through a shared context, orchestration layer, or tool output and are then accepted by another agent as if they were trusted input.
That makes the term especially important in multi-agent systems, where one compromised component can influence downstream agents that never directly saw the original attacker content. The security problem is therefore less about isolated prompt quality and more about trust propagation across agent boundaries.
Why It Becomes More Dangerous in Multi-Agent Systems
Multi-agent designs increase efficiency by letting specialised agents hand work to one another, but they also expand the trust boundary. If one agent can write to shared memory, a task queue, a message bus, or an orchestration channel, injected instructions may outlive the original interaction and be replayed in later steps.
This is why cross-agent prompt injection often behaves like a supply-chain problem inside the workflow itself. The attacker does not need every agent to be directly exposed, only one path that can seed malicious context into a downstream decision point.
In practice, the danger grows when agents can summarise, transform, or enrich each other’s outputs without strong provenance checks. A poisoned result may be rephrased as if it were a legitimate instruction, making the harmful content harder to spot than a direct prompt attack.
How Trust Spreads Across Shared Context and Tools
Cross-agent injection usually lands through one of three places: shared memory, tool output, or orchestration logic. The risky pattern is the same in each case, because one agent treats another agent’s artefact as authoritative input rather than as untrusted data.
This is why multi-agent security guidance increasingly focuses on containment, signed or scoped messages, and explicit trust boundaries. NHIMG’s Multi-Agent and A2A Security Guide is useful here because it treats agent-to-agent communication, delegation, and containment as first-class security problems. For a broader threat model of agent behaviour, Agentic AI Security Guide shows how prompt injection fits alongside tool misuse, orchestration risk, and identity concerns.
The same trust-transfer issue can also appear when one agent is effectively operating on behalf of a user or another system. In that case, a malicious instruction can be amplified by the receiving agent’s permissions, turning a content attack into an action attack.
What Defenders Need to Understand About Containment
Cross-agent prompt injection is best understood as a control failure around provenance and authority, not just content safety. The key defensive question is whether the receiving agent can distinguish trusted instructions, derived data, and attacker-controlled text before acting on it.
That distinction matters because a tool call, message, or summary may be technically correct and still operationally unsafe if it carries embedded instructions. The problem is especially visible in systems that automate handoffs across browser agents, coding agents, or customer-facing assistants, where the next step may be taken with greater privilege than the original message deserved.
For that reason, practical containment depends on narrowing what one agent is allowed to influence in another agent’s decision surface. The more a system reuses unverified context, the easier it becomes for injected content to move laterally through the workflow.
Risk and Threat Considerations
Cross-agent prompt injection creates a material risk of downstream compromise because a poisoned instruction can be reused by multiple agents, increasing the chance of data leakage, unsafe actions, or privilege abuse. The threat is strongest where agents share memory, delegate tasks, or consume each other’s outputs without provenance checks.
Failure mechanism: An attacker seeds untrusted text into one agent’s input path, then that agent forwards, summarises, or transforms the content in a way that causes another agent to treat it as trusted instruction.
Impact: The receiving agent may exfiltrate data, trigger unauthorised tool use, or propagate the malicious instruction further, expanding both blast radius and dwell time across the agent workflow.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI07 — Insecure Inter-Agent Communication | Directly covers malicious instructions moving between agents and orchestration paths. |
| ASI03 — Identity & Privilege Abuse | Cross-agent injection becomes dangerous when one agent can act with another's authority. | |
| Recommendation — Separate agent trust boundaries and validate every inter-agent message before reuse. Limit delegated authority so injected context cannot trigger higher-privilege actions. | ||
| MITRE ATT&CK | T1204 — User Execution | Covers attacker-delivered content that causes a target to execute attacker-influenced actions. |
| Recommendation — Hunt for content paths that induce agent actions and block unsafe execution triggers. | ||
| NIST SP 800-53 Rev 5 | SI-10 — Information Input Validation | Applies because agent inputs and tool outputs must be validated before being trusted. |
| AC-6 — Least Privilege | Limits the impact when injected instructions reach an agent with execution authority. | |
| Recommendation — Validate agent-fed inputs and outputs before they are passed into downstream automation. Constrain each agent to the minimum access needed for its task. | ||
Practitioner Guidance
Why practitioners should care: Cross-agent prompt injection is a systems problem, not just a prompt-hygiene problem. If agents can hand work to one another, security must account for how instructions, context, and tool output are labelled, validated, and bounded across those handoffs.
What to watch for: Treat any shared context channel, summarisation step, or delegated tool result as an untrusted boundary unless the design explicitly preserves provenance. A useful mindset is to ask whether the next agent is acting on a recommendation or on an instruction that an attacker may have smuggled into the workflow.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 10, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org