Join our Newsletter — 33% off our NHI Course

What are the signs that an agent tool catalogue is not working well in practice?

Common signs include wrong tool selection, repeated retry loops, hallucinated parameters, and rising token usage from oversized schemas. The article also points to missing functional descriptions and poor failure guidance as frequent causes. When these symptoms appear, the catalogue is too API-shaped and too broad for reliable agent execution.

What the symptoms are telling you about the catalogue

When an agent tool catalogue is healthy, the agent can choose the right tool from a small, well-described set and execute it with predictable parameters. When it is not working well, the symptoms are usually operational, not cosmetic: the agent hesitates, chooses poorly, retries the same action, or starts improvising around the catalogue instead of using it cleanly. That is a sign the catalogue is not giving the model enough semantic guidance to map intent to tool.

A catalogue can fail even when every tool is technically reachable. The problem is often that the entries are too API-shaped, too broad, or too similar, so the agent cannot tell which tool is intended for which job. In practice, that shows up as tool confusion, brittle selection, and a widening gap between the user’s request and the action the agent actually takes.

The most reliable clue is repetition. If the agent repeatedly asks for the same missing fields, cycles through similar tools, or keeps attempting calls that should have been disambiguated by the catalogue, the catalogue is not acting like a decision aid. It is acting like an unhelpful inventory.

How to recognise poor tool design in execution traces

Execution traces usually reveal the pattern faster than a design review. Wrong tool selection, repeated retry loops, and hallucinated parameters indicate the agent is not learning enough from the catalogue entry to complete the action. If the agent also consumes excessive tokens to read oversized schemas, the catalogue is forcing the model to spend effort parsing structure instead of deciding.

That failure pattern often points to missing functional descriptions. A tool name and a request schema are rarely enough for an agent to choose well. The catalogue needs to explain purpose, expected inputs, failure conditions, and the difference between near-duplicate tools. Without that context, the model compensates by guessing, overcalling, or trying several tools in sequence.

Another warning sign is poor failure guidance. If a tool returns an error but the catalogue does not tell the agent what to do next, the agent may simply retry the same path, change parameters at random, or switch to a worse alternative. Good catalogues help the model recover from failed execution, not just start it.

Why broad, API-shaped catalogues break agent behaviour

A catalogue becomes fragile when it mirrors an underlying API too closely. Human developers may be comfortable choosing from many endpoints, but an agent needs a sharper abstraction layer. If every endpoint is exposed as a tool, the agent gets a large, flat choice set with weak distinctions, which increases selection error and makes prompt interpretation harder.

Overbreadth also increases ambiguity across similar actions. If several tools look interchangeable, the agent may pick the wrong one based on wording rather than intent. If the schemas are oversized, the agent may also use the right tool badly, because it spends tokens on irrelevant fields and loses signal about the few fields that actually matter.

For that reason, catalogue quality is not just about coverage. It is about curation: grouping related actions, naming them by business intent, and making the success path obvious. The best catalogues reduce the amount of reasoning the agent must do before execution.

Risk and Threat Considerations

Poor tool catalogues create reliability risk first, but they also create security exposure when the agent can be nudged into the wrong action or the wrong scope. Tool confusion, retry behaviour, and overly broad schemas can increase the chance of unintended side effects, especially when a tool can read, modify, or forward sensitive data.

Failure mechanism: The agent cannot distinguish similar tools or recover cleanly from failed calls, so it retries, substitutes parameters, or selects a broader tool than intended. That creates avoidable execution drift and can expand the blast radius of a mistaken action.

Impact: Operators see higher failure rates, higher token usage, and less trustworthy automation. In more sensitive workflows, poor catalogue design can also lead to accidental access, overbroad operations, or hard-to-audit actions that are difficult to attribute after the fact.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
OWASP Agentic AI Top 10 ASI02 — Tool Misuse Agent tool selection errors and retry loops are direct tool-misuse symptoms.
ASI03 — Identity & Privilege Abuse Poor catalogue design can lead agents to take broader actions than intended.
ASI08 — Cascading Failures Retry loops and oversized catalogs can compound into repeated execution failures.
Recommendation — Constrain tool choice with clearer intent boundaries and fail-closed recovery paths. Limit per-tool authority so mistaken selection cannot expand privilege or impact. Reduce tool coupling and define safe fallback behaviour to stop failure cascades.
NIST SP 800-53 Rev 5 AU-2 — Event Logging Trace review is needed to spot wrong tool use, retries, and parameter drift.
SC-3 — Security Function Isolation Separating tools by purpose reduces ambiguous execution paths for agents.
Recommendation — Log tool selections, parameter sets, and failures for agent execution review. Isolate high-impact actions into narrower tools with distinct purposes.

Practitioner Guidance

What to verify: Check whether each tool entry answers four questions for the agent: what the tool does, when to use it, what a success looks like, and what to do when it fails. If any of those are missing, expect brittle selection and noisy retries.

Decision rule: If two tools can plausibly satisfy the same intent, do not rely on schema detail alone to separate them. Add an intent-level description, narrow the tool surface, or split the catalogue by task family so the agent has a clearer choice.

What good looks like: The agent chooses the intended tool on the first or second attempt, uses only the needed parameters, and fails over to a clearly stated recovery path without cycling through near-duplicates. That is a stronger signal than raw tool count or catalogue size.

Practitioner takeaway: A tool catalogue is working well when it reduces decision burden for the agent, not when it exposes the most endpoints or the most schema detail.