Common warning signs include an agent acting on public issues as if they were trusted tasks, unexpected references to private repository content, and output that includes credentials, environment variables, or other sensitive snippets. Another red flag is when a workflow can move from reading user-submitted text to committing changes or generating documentation without a review step. Those are signs that context boundaries are too weak.
How to read the failure pattern in an MCP workflow
An MCP workflow is failing to contain untrusted input when the boundary between user-originated content and trusted execution has become porous. That usually shows up as a workflow treating low-trust text as if it were approved context, reusing it across tools or steps without a check, or letting it shape actions that should have stayed gated. The key question is whether the input can influence behaviour outside its intended trust level.
That boundary failure is often visible before anything obvious breaks. A workflow may seem functional while quietly propagating untrusted context into retrieval, summarisation, file writes, or tool calls. Once that happens, the issue is not just content quality, it is control failure: the system is no longer reliably separating read-only inputs from decision-making, execution, or privileged side effects.
What the warning signs usually look like
The strongest indicator is context collapse, where the workflow starts to act on user-submitted content as though it came from an internal instruction source. A second sign is leakage across trust zones, such as private repository details, hidden system notes, or secret-like material appearing in outputs after the workflow handled external text. When the workflow begins to reuse the same context for interpretation and execution, containment is already weak.
Another practical warning is step skipping. If the workflow can move from ingesting untrusted text to creating tickets, editing files, or generating documentation without a review or approval boundary, then untrusted input is participating in a trusted action path. MCP Security Guide is useful background for understanding why authorization boundaries and token handling matter when the protocol connects external input to tools.
You should also watch for outputs that contain credentials, environment variables, API keys, or other sensitive snippets that were not explicitly supplied for that purpose. That is not just a leakage problem, it is evidence that the workflow has let untrusted material interact with protected context or with a tool that had excess access. At that point, the workflow is not containing input, it is amplifying it.
Where containment usually breaks, and why it matters
Containment breaks most often at trust transition points: ingestion, retrieval, tool selection, and write-back. If the workflow does not clearly distinguish between text that can be read, text that can influence a model, and text that can trigger an action, then untrusted input can move too far through the chain. In MCP-heavy systems, that risk is magnified when a client, server, or downstream tool accepts context without validating whether it should be trusted at all.
External guidance on agentic security is helpful here because the failure mode is not unique to one protocol. OWASP Agentic AI Top 10 frames the broader class of problems that appear when agents over-trust input, misuse tools, or inherit privilege they should not have. Model Context Protocol: Authorization specification is also relevant because audience-bound tokens and no token passthrough are exactly the kinds of design choices that limit uncontrolled propagation.
The operational consequence is simple: once untrusted input can steer a trusted action, the workflow can be tricked into disclosure, false reasoning, unauthorized file changes, or unsafe tool invocation. Even when the output looks plausible, the system may no longer be making decisions from verified context. That is why containment failures should be treated as security defects, not only prompt-engineering problems.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | MCP workflows fail when untrusted input gains trusted agent authority. |
| ASI02 — Tool Misuse | The failure mode includes unsafe tool invocation driven by untrusted context. | |
| Recommendation — Bound tool and action authority so untrusted input cannot inherit privileged agent behavior. Require explicit gating before tools act on user-supplied or retrieved content. | ||
| OWASP API Security Top 10 | API5 — Broken Function Level Authorization | Untrusted input causing unintended actions reflects broken action-level control. |
| Recommendation — Enforce function-level authorization on every side-effecting workflow step. | ||
| NIST SP 800-53 Rev 5 | AU-6 — Audit Review, Analysis, and Reporting | Containment failures are easier to confirm when tool use and sensitive outputs are logged. |
| AC-6 — Least Privilege | If untrusted context can influence execution, excess privilege increases blast radius. | |
| Recommendation — Review logs for boundary crossings, sensitive output, and unauthorized tool use. Reduce tool and file-system privileges so untrusted input cannot trigger broad impact. | ||
Practitioner Guidance
What to verify: Check whether the workflow has explicit trust boundaries between ingestion, reasoning, retrieval, and execution. If the same context can both interpret user text and drive side effects, assume containment is weak until proven otherwise.
Decision rule: If untrusted input can reach a tool, file write, or privileged action without a separate approval or validation step, treat the workflow as unsafe even if no incident has occurred yet. The absence of visible abuse does not mean the boundary is holding.
What good looks like: Untrusted content can be analysed, but it cannot silently upgrade itself into instructions, credentials, or execution authority. Review points are explicit, sensitive material stays out of generated output, and tool use is bounded by the minimum necessary context.
Practitioner takeaway: The test is not whether the workflow can process untrusted input, it is whether that input can cross into trusted execution without being stopped, sanitised, or re-authorised first.
Related resources from NHI Mgmt Group
- What are the signs that a dashboard plugin may be failing to handle untrusted input safely?
- What are the risks of using static credentials in MCP servers?
- What fails when an automation workflow can pass untrusted input into a mail node?
- What are the signs that an MCP authorization flow is failing in practice?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org