Common signs include the agent ignoring a tool that clearly fits the request, choosing a generic or unrelated tool, or failing when parameters are incomplete. Vague names, unclear descriptions, and deeply nested schemas all reduce match quality. If the tool works in code but not in natural language, the description is usually the problem.
When an MCP tool description is too weak, what actually goes wrong?
An mcp tool description is too weak when the model cannot reliably tell what the tool is for, what inputs it expects, or when it is better than a competing option. The result is not just missed matches, but degraded tool arbitration: the agent hesitates, misroutes the request, or treats a broadly named utility as interchangeable with a specific one.
Weak descriptions usually fail at the selection layer before execution even starts. A model may still “know” a tool exists, but if the description does not clearly express intent, capability boundaries, and decision cues, it cannot map a natural-language request to the right action with confidence.
In practice, the problem is often semantic rather than technical. The tool may be callable, correctly implemented, and usable in code, yet still underperform in agent selection because the description does not speak the model’s matching language: task intent, parameter expectations, output shape, and operational context.
Which description signals are most likely causing selection failure?
The strongest signals are vague names, description text that is too generic to distinguish one tool from another, and missing cues about parameters or constraints. If the agent cannot infer MCP authorization and tool-bound context from the description, it may default to a safer or more generic path, even when the intended tool is the best fit.
Deeply nested schemas can also weaken selection because the model has to infer too much structure before it can decide whether the tool is appropriate. That creates a poor match between natural-language intent and the tool’s actual affordances, especially when the description does not explain the schema in plain terms.
Another warning sign is when the tool is obvious to a human reviewer but not to the model. That usually means the description relies on internal jargon, implementation detail, or implied knowledge instead of stating the practical job the tool performs and the conditions under which it should be chosen.
Tool descriptions also fail when they blur scope. If multiple tools claim overlapping purposes without clear differentiators, the model tends to pick the first plausible one or the most generic one. Good descriptions reduce that ambiguity by making the decision boundary explicit.
How can you tell the description, not the code, is the real problem?
The clearest test is the code-versus-language mismatch: if a tool works reliably when invoked directly, but the agent still ignores it or mis-selects it from a natural-language prompt, the description is usually underspecified. That points to a retrieval and ranking problem, not an execution bug.
Another sign is repeated failure on requests that should be straightforward from the tool’s actual behavior. If the model only selects the tool after the prompt is rewritten with the exact schema vocabulary, the description is not surfacing the right decision cues for normal user language.
Selection problems also show up when the model can use the tool after being explicitly reminded of it, but does not reach for it on its own. That means the tool is visible in the catalog but not salient enough in the description for autonomous selection.
For agentic tool ecosystems, this is why strong descriptions matter more than feature lists. A description should help the model answer three questions quickly: what does this tool do, what kind of request should trigger it, and what makes it different from nearby tools?
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack and risk surface, while CIS Controls v8 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | MCP tool selection failures stem from poor tool choice and misuse signals. |
| ASI03 — Identity & Privilege Abuse | Tool choice is tied to delegated authority and action boundaries for agents. | |
| ASI07 — Insecure Inter-Agent Communication | Weak descriptions can blur how tools should be invoked across agentic workflows. | |
| Recommendation — Strengthen tool descriptions so the agent selects the intended tool and avoids generic fallbacks. Define tool purpose and boundaries so agents do not route actions to overbroad capabilities. Make invocation intent explicit so downstream agents can select and call tools correctly. | ||
| OWASP API Security Top 10 | API5 — Broken Function Level Authorization | Tool descriptions influence which functions an agent reaches and uses. |
| Recommendation — Document function boundaries clearly so agents do not invoke the wrong capability set. | ||
| CIS Controls v8 | CIS-16 — Application Software Security | Tool descriptions are part of secure software behavior when agents consume them directly. |
| Recommendation — Treat tool metadata as a security-sensitive interface and keep it precise and testable. | ||
Practitioner Guidance
What to verify: Check whether the description contains the task intent, the expected input shape, and a clear boundary against adjacent tools. If a reviewer has to read the schema to understand the purpose, the description is probably too weak for reliable agent selection.
What good looks like: The model should pick the tool on the first pass for ordinary user phrasing, without needing schema vocabulary or manual prompt steering. Good descriptions make the right tool feel discoverable, not merely callable.
Common mistake: Teams often describe tools from an implementation viewpoint, then assume the agent will infer the use case. In practice, the model needs decision-ready language, not just functional accuracy.
Practitioner takeaway: If a tool is technically sound but inconsistently selected, treat the description as part of the interface surface, not documentation; the model can only choose what the text makes distinguishable.
Related resources from NHI Mgmt Group
- What are the signs that AI agent governance is too weak for production use?
- What are the signs that alert grouping is too weak to support effective investigation?
- What are the signs that AI agent telemetry is too weak for investigation?
- What are the signs that AI agent security controls are too weak?