A one-time review fails when a server can change behavior after it is trusted. A malicious MCP server may look harmless during initial checks, then rewrite tool descriptions or return instructions only after several calls. Security teams should monitor tool metadata for drift after approval and compare cached definitions on every reconnect, because the attack can begin only once the server is already in use.
What fails when MCP metadata is trusted only once?
A one-time review assumes the server stays honest after approval, but mcp server can change what they advertise or return later. That means initial checks can miss tool-description drift, delayed instruction injection, or other post-approval behavior changes. The security problem is not just the first trust decision, it is whether the metadata remains stable after the server starts operating.
When metadata is treated as static, the approval process becomes a point-in-time snapshot instead of an ongoing control. In practice, that creates a gap between what was reviewed and what the client actually executes later, especially when the server can alter tool names, descriptions, or responses after a few interactions.
Practitioners should treat metadata as part of the attack surface, not just onboarding paperwork. The relevant question is whether the client can detect that a previously approved MCP server now presents different tool definitions, different intent, or different behavior than it did during review.
Why post-approval metadata drift is more dangerous than a bad first impression
The failure mode is delayed trust abuse. A server can appear well-behaved long enough to pass an initial assessment, then alter the contents that agents rely on for tool selection and execution. That is especially risky when the client caches definitions, reuses previous judgments, or assumes a trusted server will remain consistent across sessions.
This pattern matters because the decision boundary sits before the abuse. If review happens only once, the organization may never see the malicious state that emerges later. That is why the strongest control is continuous comparison of the current metadata against a trusted baseline, not a single approval event.
For MCP-specific control design, the MCP authorization specification is relevant because it frames the server as an OAuth-protected resource with explicit authorization expectations. The protocol point is not only who may connect, but whether the server’s published behavior remains consistent enough to trust after connection.
What the attack looks like in practice
The attacker objective is to gain execution influence after the server is already trusted. One common path is to delay the malicious change until several calls have occurred, because the early interaction establishes confidence and may seed cached definitions in the client or agent runtime.
Once the metadata shifts, the agent can be steered through rewritten tool descriptions, altered responses, or instructions that appear to come from an approved source. That is why this issue is closely related to tool poisoning and instruction smuggling, even when the initial approval looked clean.
For broader agentic security context, OWASP Agentic AI Top 10 captures the wider class of risks around tool misuse, identity and privilege abuse, and trust exploitation. For the protocol side, RFC 9728 is the underlying metadata model that helps explain why protected resource metadata must be handled carefully and not treated as immutable by assumption.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST SP 800-53 Rev 5 and OWASP ASVS set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | MCP metadata drift can steer agents into unsafe tool use. |
| ASI09 — Human-Agent Trust Exploitation | One-time review can create misplaced trust in a later-changed server. | |
| ASI03 — Identity & Privilege Abuse | A trusted MCP server may later exploit its accepted authority. | |
| Recommendation — Validate live tool metadata before allowing agent tool execution. Require ongoing trust checks for approved agent-connected services. Reassess server authority when metadata or behavior changes. | ||
| OWASP Non-Human Identity Top 10 | NHI-06 — Insecure Cloud Deployment Configurations | Post-approval metadata drift is a configuration-control failure surface for MCP servers. |
| NHI-10 — Human Use of NHI | Review-only approval can leave humans trusting stale server behavior. | |
| Recommendation — Continuously compare approved server metadata against the live configuration. Keep human approval tied to live verification, not a one-time signoff. | ||
| NIST SP 800-53 Rev 5 | SI-4 — System Monitoring | Drift detection requires ongoing monitoring of server behavior and metadata. |
| CM-3 — Configuration Change Control | Metadata changes after approval need controlled review and authorization. | |
| AU-6 — Audit Record Review, Analysis, and Reporting | Comparing cached and live definitions depends on reviewable records of change. | |
| Recommendation — Monitor live MCP metadata for unauthorized or unexpected changes. Treat tool metadata changes as controlled configuration changes. Review audit evidence for metadata changes and reconnect mismatches. | ||
| OWASP ASVS | V15 — Secure Architecture | The answer concerns trust boundaries and runtime behavior changes in an agent-facing service. |
| Recommendation — Design agent integrations to validate server behavior at runtime, not once. | ||
| OWASP API Security Top 10 | API8 — Security Misconfiguration | Mutable tool metadata is a misconfiguration-style exposure when clients trust stale definitions. |
| Recommendation — Reject stale capability data and refresh trusted definitions on reconnect. | ||
Practitioner Guidance
What to verify: Compare current tool metadata against a known-good baseline on every reconnect and after any meaningful session change. If the server’s definitions, descriptions, or advertised capabilities change unexpectedly, treat that as a security event rather than a harmless update.
What to measure: Track metadata drift frequency, reconnect comparisons that fail, and any mismatch between cached tool definitions and live server responses. If your client cannot detect drift, you do not actually have post-approval assurance.
Decision rule: If a server can materially change the meaning of a tool after approval, do not rely on approval alone. Require runtime validation, cached-definition refresh, and a response path for sudden capability changes before the server is allowed to influence agent actions again.
Practitioner takeaway: The real control is not “did it pass review?”, it is “can we prove it still matches what was approved when the agent uses it later?”
Related resources from NHI Mgmt Group
- What breaks when an MCP server does not expose protected resource metadata?
- What breaks when MCP server metadata and authentication are not standardised?
- What breaks when MCP server UI code is not properly sandboxed and reviewed?
- What breaks when organisations treat MCP server approval as a simple yes or no decision?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org