Join our Newsletter — 33% off our NHI Course
Home› FAQ› AI Security› What happens when an MCP server returns untrusted…
AI Security

What happens when an MCP server returns untrusted content without treating it as prompt injection?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated September 30, 2026 Domain: AI Security

A maliciously crafted response can steer a review agent or similar workflow into using its own legitimate privileges against the attacker’s goals. In the reported Azure DevOps case, indirect prompt injection in pull request content led the agent to run cross-project pipelines and exfiltrate wiki data. The failure is not in authentication alone, but in trusting returned content too much.

What goes wrong when returned content is trusted too early

When an mcp server returns content that is treated as ordinary data instead of potentially hostile instruction, the client can blur the line between output and input. That is the core failure mode behind indirect prompt injection: the agent consumes the returned text, then follows it as if it were safe context. In an MCP workflow, that can turn a benign retrieval step into an instruction channel.

The practical problem is not just “bad text,” but context contamination. Once untrusted content is merged into the agent’s working context, the model may reframe attacker-supplied instructions as part of the user’s task, especially when the workflow is designed to act autonomously. That is why the risk shows up in review agents, assistants, and tool-using workflows that are allowed to act on fetched content without a trust boundary.

Why MCP makes this failure path more dangerous

MCP is designed to let tools expose structured capabilities to clients, which is useful only if the client preserves a strict separation between trusted prompts, untrusted tool output, and executable actions. If the server response can influence planning, tool selection, or downstream automation without being sandboxed, the server becomes a channel for malicious instruction. The architecture is especially sensitive when the agent holds broad access or can chain multiple tools after reading the response. MCP Security Guide explains the OAuth model, token handling, and tool-poisoning controls that matter here.

This is why “the server authenticated correctly” is not a sufficient safety claim. Authentication tells you who the server is, not whether the content it returns is safe to execute or act upon. In the Azure DevOps style failure, the dangerous step was not unauthorized login, but a trusted agent using its own legitimate permissions against the attacker’s objectives after ingesting hostile content.

That pattern is broader than one product or one incident. Any tool response that can shape agent behavior, hidden instructions in documents, code reviews, tickets, or web pages can become a prompt injection vector. Agentic AI Security Guide is useful background for the wider attack surface, including prompt injection, tool misuse, and agent guardrails.

How practitioners should contain the blast radius

The right control objective is to keep untrusted content observable, but not authoritative. Design the workflow so the agent can summarize, classify, or flag the content, yet cannot treat returned text as instructions unless a separate trust check passes. For MCP-based systems, that usually means constraining tool scope, limiting what the agent can do after retrieval, and making sensitive actions require a second, explicit decision boundary.

When the workflow includes code execution, pipeline triggers, or cross-project access, the containment bar needs to be higher. A review assistant that can launch builds or reach other repositories should be assumed to be a high-value target for indirect prompt injection. Red Teaming AI Agents for Identity Abuse helps practitioners test privilege misuse, delegation abuse, and exfiltration paths that appear only after the model is influenced by hostile content.

Sentry MCP Agentjacking 2026 is a good example of why tool-output trust must be treated as a security decision, not just a model-quality issue. If the agent can be steered by an untrusted return value, then the real control is not just authentication, it is limiting the actions that any single response can authorize.

Risk and Threat Considerations

Untrusted MCP output can become an attacker-controlled instruction stream, which creates a classic confused-deputy problem: the agent performs harmful actions using valid access, because it was induced to believe the attacker’s content was part of the task. The main risks are cross-project access, data exfiltration, unauthorized pipeline activity, and hidden persistence inside normal workflow automation.

Failure mechanism: The server’s response is merged into the agent’s context without strong trust segregation, so indirect prompt injection can redirect planning, tool use, or action selection while the agent still appears to be operating normally.

Impact: Attackers can trigger legitimate permissions for illegitimate goals, which can expose code, secrets, wiki content, tickets, or build systems and can make abuse hard to distinguish from routine automation.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10, OWASP API Security Top 10 and MITRE ATT&CK define the specific risk controls and attack patterns relevant to this topic.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseIndirect prompt injection can steer an agent into misusing its own authority.
ASI02 — Tool MisuseThe issue is hostile content driving unsafe tool selection or invocation.
ASI01 — Agent Goal HijackUntrusted output can redirect the agent away from the user’s intent.
Recommendation — Limit agent privileges and require policy checks before any high-impact action. Constrain tool access and validate tool calls against task scope. Separate untrusted content from control instructions and preserve task boundaries.
OWASP API Security Top 10API10 — Unsafe Consumption of APIsA server response is consumed unsafely when it can influence downstream behavior.
Recommendation — Treat API output as untrusted input and gate any actions it may trigger.
MITRE ATT&CKT1204 — User ExecutionThe attacker relies on content influencing a trusted actor to carry out actions.
Recommendation — Hunt for adversary content that induces legitimate users or agents to act.

Practitioner Guidance

What to verify: Verify that tool output is treated as data by default, not as executable instruction, and that downstream actions require a separate trust decision or policy gate. If a response can influence a build, deployment, or cross-repo action, treat that path as privileged.

Decision rule: If the agent can both consume external content and take real actions, assume indirect prompt injection is possible and reduce the action set before you try to “sanitize” the text. Sanitization alone does not solve a trust-boundary failure.

What good looks like: The agent can read and summarize untrusted content, but it cannot let that content alter authorization, escalation, or task scope without a deliberate human or policy checkpoint. That is the observable sign that the workflow is bounded rather than merely monitored.

Practitioner takeaway: The security question is not whether the content was authenticated, but whether it was allowed to influence authority; if the answer is yes, the agent’s own privileges become the attack path.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 30, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org