The risk shifts from model integrity alone to the trustworthiness of every upstream component that influences the agent. Compromised prompts, extensions, or MCP servers can steer behaviour without breaking the model itself. Teams should treat these inputs as part of the agent trust boundary and govern them with the same discipline used for privileged software dependencies.
How the supply chain changes when prompts, extensions, and MCP servers shape agent behaviour
The supply chain risk broadens from software integrity to behavioural integrity. A compromised prompt, extension, or MCP server may not alter the model weights at all, yet it can still redirect decisions, tool use, and data flow. That means the practical question is not only “is the model trusted?” but “which upstream components can steer the agent and with what authority?”
When those components sit in the agent path, they become security-relevant dependencies rather than convenience layers. A malicious or tampered input can induce unsafe actions, leak context, or widen access without ever defeating the underlying model.
What becomes part of the agent trust boundary
Prompts, extensions, and MCP servers influence behaviour in different ways, but they share one security property: they can change what the agent sees, what it is allowed to call, and how it interprets instructions. That makes them part of the operational trust boundary for agentic systems, much like libraries, plugins, or build inputs in a traditional software supply chain.
For MCP specifically, the trust boundary often includes server discovery, authorization, tool exposure, and token handling. MCP Security Guide is useful here because it treats the protocol as an authorization surface, not just an integration layer. For prompt and extension risk, the same logic applies: if the component can influence the next action, it can influence the security outcome.
That is why “safe model, unsafe environment” is a real pattern. The model can be intact while the surrounding control plane is compromised, misconfigured, or overly trusted.
Why behavioural supply chain risk is harder to see
Traditional supply chain checks often focus on code provenance, signatures, package integrity, or dependency CVEs. Those still matter, but they are no longer enough when the abuse path is instruction-level or tool-level rather than binary-level. A prompt injection, malicious extension update, or poisoned MCP response can steer the agent through legitimate interfaces and remain invisible to controls that only inspect the model artifact.
This is especially important for extensions and hosted tooling, because they can arrive through the same trusted channels that users rely on for productivity. Secrets in VS Code extensions 2025 illustrates how an apparently normal extension ecosystem can expose tokens and become a delivery path for broader compromise. The lesson carries over to agent ecosystems: trust in the delivery channel does not equal trust in behaviour.
MCP servers add a further twist because they can act as both capability brokers and policy decision points. If a server can enumerate tools, pass tokens, or shape the arguments the agent sends, then compromise can produce privilege misuse rather than obvious malware.
What should change in governance and review
Teams should govern prompts, extensions, and MCP servers as controlled dependencies with explicit owners, review criteria, and revocation paths. The decision point is whether the component can influence agent behaviour, access, or delegated authority. If it can, it needs a security review comparable to the review applied to privileged software dependencies.
That review should prioritise provenance, update path, permissions, and blast radius. AI Supply Chain Security and AI-BOM Guide is a useful companion because it frames models, tools, and MCP servers as supply chain assets that belong in inventory and control tracking. For MCP deployments, Model Context Protocol: Authorization specification is the key external reference for avoiding token passthrough and keeping authorization audience-bound.
Once behaviour-shaping inputs are treated as dependencies, governance becomes more concrete: inventory them, restrict who can publish or modify them, test how they change tool invocation, and remove anything that can silently expand agent authority.
Risk and Threat Considerations
Behaviour-shaping components create a supply chain where compromise can produce safe-looking but unsafe actions. The attacker does not need to break the model, only to influence the instruction, extension, or server path that the agent trusts.
Failure mechanism: A malicious prompt, extension, or MCP server injects instructions, alters tool selection, or manipulates token handling so the agent performs unintended actions through legitimate channels.
Impact: The likely outcomes are data leakage, overbroad tool use, unauthorized actions, and lateral exposure through trusted integrations, often without obvious model corruption.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and OWASP API Security Top 10 define the specific risk controls and attack patterns relevant to this topic.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI04 — Agentic Supply Chain Vulnerabilities | Prompts, extensions and MCP servers shape agent behavior through the agent supply chain. |
| ASI03 — Identity & Privilege Abuse | Compromised inputs can steer agents into excessive or misused privilege. | |
| Recommendation — Assess and restrict upstream agent inputs that can alter behavior, tools, or delegated authority. Constrain agent authority so malicious inputs cannot expand access or actions. | ||
| OWASP Non-Human Identity Top 10 | NHI-03 — Vulnerable Third-Party NHI | Extensions and MCP servers can function as third-party dependencies with exploitable trust. |
| NHI-06 — Insecure Cloud Deployment Configurations | MCP servers and hosted extensions often fail through exposed or weak deployment settings. | |
| Recommendation — Review third-party agent dependencies for insecure behavior, update risk, and trust abuse. Harden server and extension deployments so exposed configuration cannot redirect agent behavior. | ||
| OWASP API Security Top 10 | API8 — Security Misconfiguration | MCP servers expose callable interfaces whose misconfiguration can widen agent trust and access. |
| Recommendation — Lock down MCP and extension interfaces so configuration does not expose unsafe capabilities. | ||
Practitioner Guidance
What to verify: Verify that every prompt source, extension, and MCP server has an owner, a change path, and a documented purpose. If you cannot explain why the agent needs that component, it is probably carrying more trust than it should.
Decision rule: If a component can change tool calls or outputs in production, treat it as security-relevant supply chain material, not optional configuration. If it can also pass credentials or broaden scope, require explicit approval and tighter containment.
What good looks like: The agent can only consume trusted inputs from known locations, tool permissions are narrow, and any update to prompts or servers is reviewable, reversible, and attributable.
Practitioner takeaway: The key shift is from protecting the model in isolation to protecting every upstream influence that can shape what the agent does next.
Related resources from NHI Mgmt Group
- Why do AI coding agents and MCP servers increase supply chain risk on developer endpoints?
- Why do remote MCP servers reduce the risks that come with static API keys for agent access?
- What are the risks of using static credentials in MCP servers?
- Why do MCP servers change the IAM model for AI access?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org