Because an agent can keep spending tokens without failing authentication or triggering a crash. A loop, a long-running context, or injected instructions can drive repeated tool calls until budgets, rate limits, or shared quotas are exhausted. At that point, the problem is no longer only finance. It can degrade availability for legitimate agents and delay incident response work.
Why token burn becomes a security problem in agentic systems
Token consumption is not just a billing concern when it is tied to autonomous behaviour. An agent can keep looping, re-planning, retrying or following injected instructions while still appearing “healthy” from an authentication standpoint. That makes token burn a resource-control issue: it can consume shared capacity, slow legitimate work, and create an opening for abuse of trust, policy, or runtime limits.
What changes operationally is the blast radius. A single agent can turn one bad prompt, one runaway task, or one poisoned context into repeated downstream actions that affect other users, other agents, or the incident-response function itself.
How runaway usage turns into availability and control degradation
The security issue starts when token usage becomes a proxy for repeated action. If each extra model call can trigger another tool call, another retrieval, or another planning cycle, the real risk is not the token count in isolation. It is the compounded effect on compute, quotas, tool access, and shared service responsiveness. In that state, an attacker does not need to break authentication to cause harm, only to keep the agent busy.
That matters in multi-tenant or quota-driven environments because exhaustion can be uneven. A single agent may monopolize budget, saturate rate limits, or crowd out higher-priority workflows. The result is degraded availability, delayed investigation, and a weaker ability to distinguish normal autonomy from abnormal persistence.
A useful example is AI Agent Observability, Audit and Incident Response Guide, which focuses on the signals that show an agent has gone wrong and how to build a tested kill switch. The same principle applies here: if you cannot attribute rising token use to a concrete task and bounded duration, you do not really control the agent.
Why prompt injection, long context, and tool loops make it worse
Three patterns commonly turn cost into security exposure. First, prompt injection can steer an agent toward repetitive work that benefits the attacker rather than the operator. Second, long-running context can preserve the wrong goal for too long, so the agent keeps spending tokens on a stale or hostile objective. Third, tool loops can create a feedback cycle where each model turn authorizes the next action without a meaningful stop condition.
That is why AI Agent Authorisation Guide matters to this question: per-action authorization and task-scoped access are the controls that stop a model from freely converting tokens into repeated authority. It also explains why Zero Trust for AI Agents is relevant, because continuous verification and no standing privilege reduce the chance that a runaway agent can keep acting just because it is already inside the system.
When the agent is using browser sessions, shared credentials, or connected tools, the issue can escalate beyond model spend. A loop that keeps the agent alive may also keep privileged sessions alive, which increases the chance of unauthorized action, misrouting, or accidental destructive behaviour.
Risk and Threat Considerations
Unbounded token consumption can be abused as a denial-of-service path against shared AI capacity, downstream tools, or the human teams that must investigate the incident. The dangerous part is that it may look like normal workload growth until limits are already exhausted.
Failure mechanism: A malicious prompt, poisoned context, or poorly bounded task causes the agent to keep planning, retrying, or invoking tools until budgets, quotas, or response latency degrade service for legitimate work.
Impact: Availability drops, incident response slows, and the organisation may lose visibility into whether the agent is merely inefficient or actively being driven by adversarial instructions.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | Repeated tool calls and runaway action are central to token-driven abuse in agents. |
| ASI03 — Identity & Privilege Abuse | Token burn becomes a security issue when autonomy converts into excess authority or repeated action. | |
| ASI08 — Cascading Failures | Runaway usage can exhaust shared quotas and degrade dependent services and response capacity. | |
| Recommendation — Restrict tool invocation paths and enforce action-level controls to stop runaway agent execution. Bind agent actions to least privilege and require per-action authorization for sensitive steps. Limit blast radius with budgets, isolation, and failure containment for agent workloads. | ||
| NIST CSF 2.0 | PR.AA-05 — Least Privilege | Least privilege limits how much damage a spending loop can do if the agent keeps acting. |
| DE.CM-01 — Networks and systems are monitored to detect potential cybersecurity events | Rising token burn is a detectable signal of runaway or abusive agent activity. | |
| Recommendation — Apply least privilege so an agent cannot convert repeated calls into broader access. Monitor agent usage patterns and alert on abnormal sustained token consumption. | ||
Practitioner Guidance
What to verify: Check whether each agent has an explicit stop condition, a token budget, and a task boundary that is enforced outside the model. If a workflow can continue indefinitely without a human or policy checkpoint, treat that as an operational control gap, not just an efficiency problem.
Decision rule: If token growth is not explainable by a bounded task, investigate for prompt injection, runaway tool chaining, or an overly broad authorization path before you tune models or raise budgets. If the agent also has access to shared services, prioritize containment and quota isolation.
What practitioners underestimate: The main risk is often not a single expensive call, but repeated small actions that accumulate fast enough to suppress other work. The control objective is to make token spend observable, bounded, and attributable so autonomy cannot silently become a drain on security operations.
Practitioner takeaway: Treat token consumption as a security signal whenever it can drive repeated action, shared resource exhaustion, or delayed response. If spend is not bounded by policy, it is already part of your attack surface.
Related resources from NHI Mgmt Group
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org