The second agent inherits the authority of its own permissions, but not the trustworthiness of the content it receives. If the first agent browses a page, reads a tool response, or pulls plugin output and forwards it as pre-verified context, the downstream agent can execute embedded instructions with elevated access. That is the confused deputy pattern in agentic systems.
Why the receiving agent becomes a confused deputy
The key issue is not that the downstream agent lacks trust, it is that it trusts the wrong thing. Once untrusted output is presented as if it were verified context, the second agent may treat instructions, claims, or parameters inside that content as authoritative and act on them with its own permissions.
That is why this pattern is dangerous in agentic systems: the sender does not need the authority to perform the final action if it can persuade a more privileged receiver to do it on its behalf. The failure is a trust boundary failure, not a parsing bug.
When agents chain work, the receiving agent should distinguish raw observations from policy-approved facts. If that distinction is missing, a page snippet, tool response, or plugin payload can become an instruction carrier rather than a data source.
How the attack path works across agent hops
The dangerous path usually starts when one agent ingests content from a browser, connector, API, or retrieval tool and forwards it as a sanitized summary. If embedded instructions survive that handoff, the next agent may execute them because the content appears to come from a trusted upstream step.
This becomes especially risky when the first agent has broad collection rights and the second agent has broader action rights. The attacker only needs to influence the content stream once; the receiving agent then amplifies the impact by applying its own tools, accounts, or delegated access.
In practice, the same pattern appears whether the upstream source is a web page, a document, an email, or a tool response. The exploit is the same: untrusted content is mislabeled as verified context, and that label changes how the next agent handles it.
MCP Security Guide is useful here because it explains how token passthrough and tool poisoning can turn an upstream fetch into a downstream confused deputy event.
AI Agent Authorisation Guide helps frame the control question correctly: each action should be authorized for the specific agent and request, not inherited blindly from prior context.
Why this matters for containment and delegation
The practical consequence is privilege amplification. The receiving agent is acting within its own authority, so the harm can exceed what the original untrusted source could do directly. That makes delegation chains, shared sessions, and permissive handoffs especially sensitive.
Containment depends on preserving provenance and enforcing step-level checks. If an agent can consume content and also decide whether that content is safe to execute, the design has collapsed data ingestion and authority in a way attackers can exploit.
Zero Trust for AI Agents is a strong companion here because it treats each principal and request as separately verified, instead of assuming trust across the chain.
AI Agent Observability, Audit and Incident Response Guide supports the operational side of this problem by making it easier to trace which upstream content preceded a bad downstream action.
Risk and Threat Considerations
This pattern is risky because it lets attacker-controlled text cross a trust boundary and gain effect through a more privileged agent. The immediate exposure is unauthorized tool use, but the broader concern is lateral impact through delegated access, credentials, or connected systems.
Failure mechanism: A source agent forwards untrusted content as verified context, and the receiving agent executes embedded instructions or treats them as policy-approved facts.
Impact: Attackers can steer privileged actions, trigger data exposure, or induce destructive tool use without directly compromising the higher-privilege agent.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Explains privileged downstream action triggered by untrusted agent content. |
| Recommendation — Enforce per-action authorization so one agent cannot confer privilege to another. | ||
| NIST SP 800-53 Rev 5 | IA-9 — Service Identification and Authentication | Agent-to-agent handoffs require authenticated, attributable service interactions. |
| AC-6 — Least Privilege | Limits blast radius when a receiving agent is tricked into acting on bad content. | |
| Recommendation — Authenticate each agent boundary and reject unauthenticated context transfers. Restrict each agent to the minimum access needed for its own task. | ||
| NIST Zero Trust (SP 800-207) | Zero Trust Architecture | Separates trust decisions from content provenance across agent hops. |
| Recommendation — Verify every request and context handoff instead of trusting prior steps. | ||
| MITRE ATT&CK | T1204 — User Execution | Embedded instructions in content rely on the recipient to execute attacker-supplied guidance. |
| Recommendation — Detect and block instruction execution that originates from untrusted content. | ||
Practitioner Guidance
What to verify: Verify that every inter-agent handoff preserves provenance, source classification, and a clear distinction between observations and instructions. If the receiver cannot tell whether content was user-supplied, retrieved, or tool-generated, it should not treat that content as executable guidance.
Decision rule: If a downstream action could be harmful when driven by malicious content, require explicit re-validation or human approval before execution. The more privilege the receiving agent has, the stricter the handoff should be.
What good looks like: The downstream agent can cite what it saw, where it came from, and why it is allowed to act on it, while refusing to execute any embedded instruction that was not independently authorized.
Practitioner takeaway: The control objective is not to make agents distrust everything, but to ensure that untrusted content never gains authority simply because a more privileged agent relayed it.
Related resources from NHI Mgmt Group
- What happens when an agent combines untrusted content with sensitive access?
- What happens when one agent writes poisoned content into a shared memory store?
- What happens when an MCP-enabled agent reads untrusted content and executes hidden instructions?
- What happens when an AI agent passes data to another agent without a chain of custody?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org