Teams often assume a language model can safely work with developer-oriented APIs without adaptation. In practice, raw APIs expose pagination, complex schemas, and unclear error handling that increase failure rates and token waste. Agents need tools designed for structured calling, predictable outputs, and constrained actions. Without that layer, reliability falls and governance becomes much harder.
Why raw APIs are a poor default surface for AI agents
Raw APIs are built for developers, not for autonomous systems that need bounded, repeatable behaviour. An agent calling a developer-facing endpoint has to infer pagination, interpret schemas, recover from ambiguous failures, and decide what to do next. That is exactly where reliability drops: the model spends tokens on navigation instead of the task, and the tool surface becomes harder to govern.
For agents, the issue is not that APIs are “too technical”, but that they are often too open-ended. A good agent tool abstracts the API into a smaller action space with clear inputs, fixed output shapes, and explicit limits on what the agent is allowed to do. That keeps the model focused on decision-making rather than protocol handling.
There is also a control problem. A raw API usually exposes more of the underlying system than the task needs, so the agent can stumble into unnecessary endpoints, large result sets, or actions that should have been separated by policy. In practice, the right design is often a constrained tool layer that translates intent into safe, predictable API calls.
Where reliability breaks down in practice
The most common failure mode is hidden complexity. Pagination, retries, rate limits, nested schemas, and sparse error messages force the agent to maintain state across calls while still planning the next step. That is fragile even when the model is competent, because a small interpretation error can cascade into failed workflows or repeated calls that waste context and budget.
Another issue is output instability. If the API returns variable fields, long payloads, or mixed success and warning states, the agent has to infer what counts as completion. That makes it harder to verify whether the task succeeded, and harder to build a deterministic guardrail around the workflow. A structured tool wrapper reduces that uncertainty by returning only the fields the agent actually needs.
This is where API-specific guidance matters. The OWASP API Security Top 10 is useful because it frames the kinds of API exposure teams should be thinking about, especially broken authorisation, unrestricted consumption, and insecure configuration. The agentic problem often starts as a usability issue and then turns into an API security issue when the same broad endpoint can be invoked too freely.
How to design the agent layer instead of exposing the whole API
Teams get better results when they separate the agent’s reasoning from the raw interface. The agent should call small, purpose-built tools that represent a business action, not a generic endpoint. In many cases, a wrapper should validate inputs, constrain output shape, and handle pagination or retries behind the scenes so the agent sees one stable operation instead of several brittle API steps.
That design also makes governance easier. When the tool is narrow, policy can be applied per action rather than per endpoint family, and reviewers can reason about what the agent is actually permitted to do. NHIMG’s AI Agent Authorisation Guide is a useful companion here because it treats least privilege, task-scoped access, and approval gates as design requirements rather than afterthoughts.
For agent identity and control boundaries, the useful reference point is a zero trust mindset: verify the requester, constrain the action, and avoid standing privilege. NHIMG’s Zero Trust for AI Agents is relevant because raw API access tends to blur those boundaries, while a tool layer can preserve them.
For teams building the broader pattern, NHIMG’s MCP Security Guide helps show why structured tool mediation is usually safer than letting an agent improvise against arbitrary endpoints.
Risk and Threat Considerations
Raw APIs increase the blast radius of agent mistakes because the agent is interacting directly with a general-purpose interface, not a bounded task surface. That creates exposure to overbroad access, accidental data retrieval, repeated calls that amplify cost or rate-limit pressure, and action chains that are hard to audit after the fact.
Failure mechanism: The agent misreads schema, pagination, or error states, then retries, over-fetches, or invokes a broader endpoint than intended. If the API also has weak authorisation or weak consumption limits, the same design flaw can become an abuse path rather than just a reliability issue.
Impact: Teams see failed workflows, noisy automation, higher token spend, and weaker governance over what the agent actually did. In the worst case, the agent can expose sensitive data, trigger unintended changes, or create a control gap that is difficult to detect quickly.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP API Security Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP API Security Top 10 | API8 — Security Misconfiguration | Raw APIs expose broad, error-prone surfaces that agents can misuse. |
| API4 — Unrestricted Resource Consumption | Agents can waste tokens and trigger repeated calls through pagination and retries. | |
| API5 — Broken Function Level Authorization | Agent tool calls need action-level permission boundaries, not broad endpoint access. | |
| Recommendation — Constrain exposed endpoints and normalise responses before agents can call them. Add quotas and response limits to stop agent-driven overconsumption. Enforce function-level checks on every agent-invoked action. | ||
| NIST SP 800-53 Rev 5 | AC-6 — Least Privilege | Agent tool access should be minimized to the task surface. |
| IA-9 — Identification and Authentication (Service and Device Accounts) | Agent-to-API interaction depends on controlled machine authentication. | |
| Recommendation — Grant the agent only the privileges needed for the specific workflow. Use service-account authentication with tightly scoped credentials for tool calls. | ||
Practitioner Guidance
What to prioritise: Treat “raw API access” as an exception, not the default. The first design question should be whether the agent truly needs endpoint-level freedom, or whether one constrained tool can safely express the task.
What to verify: Confirm that the tool layer owns pagination, schema normalisation, retries, and error translation. If the model still has to reason about those mechanics directly, the interface is probably too raw for reliable agent use.
Common mistake: Teams often test whether the agent can call the API once, then assume the integration is ready. The better test is whether it can complete the full workflow repeatedly, with predictable output and a reviewable action trail.
Practitioner takeaway: Good agent integration is usually about reducing degrees of freedom, not increasing them. If the agent must improvise through a complex API, reliability, cost, and governance all degrade together.
Related resources from NHI Mgmt Group
- What do security teams get wrong about zero-click attacks against AI assistants and agents?
- What do security and operations teams get wrong about using raw logs with AI agents?
- What do security teams get wrong about prompt engineering for AI agents?
- What do security teams get wrong about prompt filtering for AI agents?