TL;DR: Unosecur says prompt injection in MCP becomes materially more dangerous when an AI agent can use tools, because malicious instructions can drive real actions such as data retrieval, messaging, code execution or API calls. Security now has to control execution, not just model output, because poisoned context can cross directly into privileged tool use.
At a glance
What this is: This analysis argues that MCP prompt injection is an execution-control problem because poisoned context can trigger real tool use, not just bad model output.
Why it matters: IAM and security teams need to treat tool authorization, scoped credentials and execution-layer logging as part of identity governance for AI agents and other non-human identities.
By the numbers:
- 24,008 unique secrets were exposed in MCP configuration files in 2025 alone, the protocol's first year of widespread adoption, according to the State of Secrets Sprawl 2026.
👉 Read Unosecur's analysis of prompt injection and MCP execution control
Context
Prompt injection in MCP is a governance problem because the model can move from interpreting content to invoking tools with real side effects. Once an agent can read files, query databases, send messages, execute code or call APIs, malicious instructions can influence execution rather than only produce a wrong answer.
The risk sits at the boundary between untrusted context and authorised action. In identity terms, that makes the question one of which tools, credentials and data paths a non-human identity may touch when its inputs are not trustworthy.
Key questions
Q: What breaks when an MCP-connected agent can turn untrusted text into tool actions?
A: The boundary between data and execution breaks. In an MCP workflow, poisoned content can influence later tool selection, data access or outbound requests, so the risk is not only a bad answer. Security teams need controls that constrain what the agent may do, not just what it may say.
Q: Why does prompt injection become more dangerous when a model can use tools?
A: Because the output stops being just text. A compromised instruction path can become a real action, such as deleting data, revealing secrets, or changing records. The larger the attached permissions, the larger the blast radius, so tool access must be treated like delegated privilege, not a convenience feature.
Q: What are the signs that MCP tool control is failing?
A: Watch for tool calls that follow external content retrieval, unexpected instructions embedded in tool metadata, calls to resources outside the agent's normal task, or sensitive data appearing in unrelated parameters. A second warning sign is when approval logs cannot explain why a specific action was taken.
Q: How should teams separate legitimate automation from unauthorised execution in MCP?
A: By making execution rights explicit and narrow. Read, write, delete and external transmission should not share the same authority, and high-impact actions should require approval that shows the concrete operation and target. That way, a poisoned prompt cannot automatically inherit the ability to act.
Technical breakdown
How prompt injection becomes tool execution in MCP
MCP standardises how models discover and invoke tools, which means tool metadata, schemas, parameters and return values become part of the decision surface. If attacker-controlled text reaches the model through a webpage, document, issue, email or tool response, the model may treat that text as an instruction and choose a later tool call accordingly. That is why prompt injection in this setting is not just about unsafe language output. It is about untrusted content being able to influence an execution path that can reach files, databases, messaging systems or APIs.
Practical implication: treat every tool-facing input and output as a potential execution trigger, not merely model context.
Why tool poisoning, shadowing and rug pulls matter
Tool poisoning occurs when hostile instructions are embedded in descriptions, schemas or returned content. Tool shadowing changes how a model interprets a trusted tool by placing a maliciously similar one alongside it, while a rug pull alters tool behaviour after approval. These patterns exploit the fact that MCP exposes multiple trust-bearing objects to the client, not just the final tool call. The result is cross-tool escalation, where low-trust content can steer a high-privilege action through apparently legitimate interactions.
Practical implication: govern tool registration, change control and approval semantics as identity controls, not just developer hygiene.
Why least privilege must apply at the execution layer
The core failure mode is excessive permission inheritance. A model that can summarise text should not automatically inherit write access, account-change rights or external transmission privileges because another component in the workflow needs them. MCP guidance therefore points to sandboxing, scoped credentials, structured outputs and execution-layer authorisation as controls that survive bad reasoning. The security boundary has to sit where the action happens, because system prompts cannot reliably stop a model from following malicious instructions once those instructions reach a tool call.
Practical implication: separate read, write, delete and outbound transmission permissions so a compromised prompt cannot escalate across tool classes.
Threat narrative
Attacker objective: The attacker wants trusted tool execution to carry out data theft, unauthorised actions or downstream system changes under the cover of a legitimate agent workflow.
- Entry occurs when malicious instructions are embedded in a webpage, document, email or tool response that an MCP-connected agent ingests as data.
- Credential or authority abuse follows when the model turns that poisoned context into a tool selection or an outbound request that uses granted permissions.
- Escalation happens when the agent crosses from a low-trust input into a higher-privilege tool, allowing hidden instructions to drive actions the user never intended.
- Impact is realised through real execution, such as data retrieval, messaging, code execution or API calls that expose information or change state.
Breaches seen in the wild
- Gemini CLI prompt injection flaw 2025: Tracebit showed a poisoned README could make Gemini CLI run hidden commands and exfiltrate developer secrets; Google fixed it in 0.1.14.
- Gemini AI Breach — Google Calendar Prompt Injection: Gemini AI assistant prompt injection attack leaks sensitive Google Calendar data.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
MCP prompt injection is no longer a content-safety issue; it is an execution-governance issue. The moment an agent can act on what it reads, the trust boundary shifts from model output to tool invocation. That means security teams must stop treating poisoned prompts as a nuisance and start treating them as a way to reach privileged actions through an apparently legitimate workflow. The practitioner conclusion is straightforward: govern the action path, not just the text path.
Execution-layer authorisation is the control that prompt filtering cannot replace. A system prompt can ask a model not to reveal secrets, but it cannot guarantee that a poisoned context will not steer a later tool call. The control failure here is not a missing filter, it is a misplaced assumption that reasoning safety is enough to protect execution. Practitioners need to recognise that tool permissions must be evaluated independently of model intent.
Prompt injection in MCP exposes the identity blast radius of over-shared tool privileges. When one agent component can read, write and transmit across multiple tools, a single compromised context can affect unrelated systems. That is why least privilege in MCP has to be scoped by tool class, destination and action type. The practitioner conclusion is to reduce the number of paths a poisoned context can traverse.
Indirect prompt injection turns ordinary content ingestion into a trust propagation problem. Tool descriptions, schema fields and returned content can all carry attacker-controlled instructions, so trust can move across multiple layers before a human ever sees an approval prompt. The important governance question is which inputs are allowed to influence future execution, not whether the model sounded compliant at the time. Practitioners should assume trust can be laundered through metadata.
Tool annotations and execution traces need to be treated as audit evidence, not just telemetry. If security teams cannot reconstruct which content influenced which tool call, they will misclassify malicious execution as normal API use. That weakens incident response, recertification and accountability for non-human identities. The practical conclusion is to require traceable execution records that link context, decision and action.
From our research library:
- 71% of organizations use third-party APIs, according to Gartner’s 2024 data.
- Read next: MCP Security Guide
What this signals
Execution-control gap: MCP changes the governance question from whether a model can be influenced to whether the resulting action can be contained. That is a material shift for AI agent programmes, because the relevant control point is now tool authorisation and not prompt hygiene.
Trust propagation is the hidden risk: descriptions, schemas, returned content and outbound parameters all carry trust, so one weak integration can influence a stronger one. Practitioners should map where untrusted content can still flow into higher-privilege tools, because that is where the identity blast radius expands.
24,008 unique secrets were exposed in MCP configuration files in 2025 alone, the protocol's first year of widespread adoption. That scale makes execution-layer governance urgent for teams already exposing tools through agents and MCP servers.
For practitioners
- Scope each tool by action type Assign separate permissions for read, write, delete, outbound transmission and code execution so a summarisation task cannot inherit broader authority.
- Sandbox local MCP servers Run local or self-hosted MCP servers with restricted filesystem, network and system access so a compromised server cannot operate with full client privileges.
- Require explicit execution approvals Show the exact action, target, data and destination before approving high-impact tool calls, rather than relying on a generic allow or deny prompt.
- Log the full execution chain Record tool discovery, description changes, model decisions, arguments, results, authorisation outcomes and downstream actions so poisoned context can be reconstructed later.
Key takeaways
- MCP prompt injection becomes an execution problem when poisoned context can drive tool calls, not just generate unsafe text.
- The main governance gap is over-shared privilege across tools, because one compromised input can cascade into data access or state-changing actions.
- Containment has to move to the execution layer, with scoped permissions, explicit approvals and traceable tool activity.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-04 — Insecure Authentication | Tool-poisoned MCP flows rely on weak trust boundaries around non-human execution authority. |
| NHI-05 — Overprivileged NHI | The article centres on excessive permission inheritance across tools and agents. | |
| NHI-10 — Human Use of NHI | Human approval quality determines whether high-impact agent actions are visibly authorised. | |
| Recommendation — Scope MCP tool access so untrusted context cannot trigger authenticated actions outside the agent's job. Reduce inherited tool privilege so each MCP action class has only the access it needs. Require human review of concrete agent actions before allowing sensitive MCP tool calls. | ||
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | The article describes hostile instructions steering an agent into unsafe tool use. |
| ASI03 — Identity & Privilege Abuse | MCP prompt injection becomes dangerous when the agent can abuse granted privileges. | |
| Recommendation — Constrain agent tool routing so untrusted context cannot drive unsafe tool selection. Separate read and write authority so agent privilege cannot be reused across tool classes. | ||
| MITRE ATT&CK | TA0006;TA0008 — Credential Access; Lateral Movement | The attack path moves from poisoned content into privileged actions across tools and systems. |
| Recommendation — Map suspicious tool transitions to credential access and lateral movement behaviours in your detections. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | Execution-control in MCP depends on narrowly defined permissions and authorisations. |
| Recommendation — Apply PR.AA-05 to enforce least-privilege entitlements for each MCP-connected tool. | ||
Key terms
- Prompt Injection (Agentic): An attack where malicious instructions are embedded in content that an AI agent reads, causing the agent to execute unintended actions using its own legitimate credentials. A primary vector for agent goal hijacking and identity abuse.
- MCP tool poisoning: MCP tool poisoning is the practice of hiding malicious instructions in a tool name, description, or metadata exposed by an MCP server. Because those fields are often treated as trusted configuration, the agent may adopt the attacker’s instructions during tool selection or invocation.
- Execution-Layer Authorization: A control that approves the action itself before a tool call is allowed to execute. In agentic environments, this matters more than prompt wording because the model's decision can be wrong while the surrounding control still prevents impact.
- Trust Propagation: Trust propagation is the transfer of authority, context, or assumptions from one agent or system step to the next. In multi agent environments, it can turn a single compromised input or credential into a wider incident because downstream actions inherit prior trust decisions.
What's in the full article
Unosecur's full analysis covers the operational detail this post intentionally leaves for the source:
- MCP Gateway placement and request-path enforcement details for agent-to-tool traffic
- How intent classification and scoped credentials are applied before execution
- Examples of how sensitive data is stripped from tool arguments and returned content
- Trace fields captured for tool decisions, outcomes and downstream actions
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on August 11, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org