Subscribe to the Non-Human & AI Identity Journal
Home FAQ AI Security What breaks when Skills and MCP servers are…
AI Security

What breaks when Skills and MCP servers are approved without full inventory and review?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated August 1, 2026 Domain: AI Security

You lose visibility into the content and tools that can silently change agent behaviour. A Skill or MCP server can introduce hidden instructions, risky external calls, or data exposure without looking suspicious at the user prompt level. Without inventory and review, governance cannot prove which inputs were trusted and which were operationally active.

Why This Matters for Security Teams

Approving Skills and MCP servers without a complete inventory creates a governance gap that is easy to miss and hard to unwind. The issue is not just uncontrolled functionality; it is that agent behaviour can change through trusted extensions, not only through prompt text. That makes review, approval, logging, and rollback far more difficult. Current guidance in the OWASP Agentic AI Top 10 treats tool and instruction integrity as a core risk area because hidden actions can override user intent.

For security teams, the practical loss is evidence. If a Skill or MCP server is allowed into production without review, it becomes difficult to prove what was active, what data it could access, and whether it changed the agent’s effective decision boundary. That weakens change control, incident response, and compliance narratives at the same time. It also creates a blind spot between AI governance and operational security, where owners may assume the agent is behaving as designed while the tool layer is quietly expanding its reach. In practice, many security teams encounter the damage only after an unexpected external call, data leak, or workflow change has already occurred, rather than through intentional review.

How It Works in Practice

Skills and MCP servers sit between the model and the systems it can reach, so they need the same discipline normally applied to software dependencies and privileged integrations. A full inventory should record the owner, purpose, permissions, data handled, version, update source, and approval status for each Skill or server. Review should then test what instructions it can inject, what tools it can call, and whether it can move data beyond the intended boundary. This is especially important because the model may treat tool output as trusted context unless guardrails are explicit.

Operationally, teams should treat approval as a control gate, not a one-time sign-off. That means:

  • maintaining a live register of approved Skills and MCP servers
  • validating each integration against business purpose and least privilege
  • checking for hidden prompts, unsafe defaults, and external network reach
  • requiring version review when a Skill or server changes
  • monitoring tool invocation, not just user prompts and final answers

That approach aligns well with the OWASP Top 10 for Agentic Applications 2026, which highlights risks around agent autonomy, tool misuse, and untrusted inputs. The important point is that the review scope must include both content and capability. A benign-looking Skill can still alter decision flow, and an MCP server can expose downstream systems that were never intended for that agent’s use. These controls tend to break down when teams allow dynamic plugin onboarding in fast-moving production environments because changes outpace review and ownership tracking.

Common Variations and Edge Cases

Tighter approval and inventory control often increases release friction, requiring organisations to balance speed against assurance. That tradeoff becomes sharper in environments where teams rapidly experiment with new Skills, internal MCP servers, or third-party connectors. Best practice is evolving here, and there is no universal standard for how much autonomy can be delegated before a tool becomes a governed asset rather than a convenience layer.

Edge cases appear when a Skill is read-only on paper but still influences output through retrieval, ranking, or summarisation. Another common case is an MCP server that has limited scope in design but gains broader access through inherited credentials or network paths. In both situations, the risk is not simply broken code; it is hidden authority. Where an agent operates across business units, the inventory problem becomes even more severe because approval may differ by data class, region, or workflow. For that reason, NHIMG recommends reviewing not only the tool itself but also the identity, secrets, and upstream permissions it relies on. The strongest controls are the ones that can answer a simple question after the fact: what was trusted, by whom, and for how long?

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST AI 600-1 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10A2Unreviewed Skills and MCP servers can inject unsafe instructions and tool actions.
NIST AI RMFGOVERNThis is a governance and accountability failure for agent-enabled systems.
NIST CSF 2.0PR.AC-4Approved integrations can expand access beyond intended least privilege.
NIST AI 600-1GenAI profiles emphasize controls for prompts, tools, and model outputs.
MITRE ATLASAML.TA0001Tool manipulation and hidden instructions map to AI attack techniques.

Inventory every agent tool and review it for hidden instructions, unsafe calls, and privilege expansion before approval.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on August 1, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org