Teams often underestimate how many separate consumption points need tracking. User challenges, tool executions, and worker runtime each represent different costs and governance signals. If those metrics are not visible in one place, finance, platform, and security teams can misread spend, miss quota breaches, and struggle to explain why an agent suddenly became expensive to operate.
Why AI agent usage has to be measured at more than one layer
Teams often talk about “agent usage” as if it were one metric, but the operating model usually has at least three distinct layers. A user action may trigger a request, a tool call may consume a separate quota or API cost, and a worker runtime may accumulate time, compute, or concurrency charges. Those layers need to be observable together or cost and governance drift becomes inevitable.
That separation matters because each layer answers a different operational question. User challenges show demand and approval pressure, tool executions show where the agent is spending budget and reaching external systems, and worker runtime shows how long the automation actually occupied infrastructure. If teams collapse them into one dashboard line, they can miss the point where usage turns into waste or control failure.
In practice, the most useful model is not “how much did the agent do,” but “what happened at each consumption point, and what did it cost or control?” That is especially important when an agent fans out across multiple tools or workers, because the expensive part may not be the prompt volume, it may be the repeated execution path that the prompt initiated.
Where tracking usually breaks down in real operations
The common mistake is to instrument one layer and assume it tells the whole story. For example, a platform team may capture worker CPU time but not tool-level API calls, while a finance team only sees the user-facing request count. The result is a false sense of understanding: spend looks stable, but the underlying execution pattern may be growing, fragmenting, or exceeding a quota in a different place.
Another failure mode is inconsistent attribution. If user identity, agent identity, tool identity, and worker identity are not stitched together, teams cannot explain why a single workflow became expensive, or whether the cost came from legitimate demand, retry loops, or an unexpected tool chain. That makes both budgeting and incident review slower, because the evidence is split across systems that do not share a common operational lens.
This is also where internal controls matter. A team may think it has an agent usage problem when it actually has a metering problem, because the same event is being counted once at request time, again at tool time, and again at runtime. Without a deliberate measurement model, teams can double count some activities and miss others entirely.
What good tracking looks like across finance, platform, and security
Good tracking starts by naming the unit of measure for each layer. User challenges should reflect demand and authorization flow, tool executions should reflect external action and downstream cost, and worker runtime should reflect compute occupancy and operational load. Those measures should then roll up into one view that preserves the distinctions, rather than forcing them into one blended number.
That view should also support practical decisions. Finance needs to see what drives variable cost, platform teams need to see where execution is expanding or stalling, and security teams need to see whether a worker is calling tools more often than expected or outside normal patterns. When those audiences share the same operational data, they can separate normal growth from anomalous behavior much faster.
For agent-heavy environments, a useful discipline is to measure both per-action cost and per-session or per-run cost. Per-action metrics show what a tool call or worker step consumed. Per-session metrics show whether the agent is becoming inefficient over time, such as retrying, looping, or delegating too widely. That distinction is often what explains why an agent suddenly gets expensive to operate.
Risk and Threat Considerations
When tracking is fragmented, teams can lose control of spend, miss quota breaches, and fail to notice when an agent is behaving in a way that is operationally abnormal. The same visibility gap can also hide abuse patterns, such as repeated tool invocation or runaway worker execution, until the impact is already material.
Failure mechanism: Separate metering paths are interpreted as one usage story, so the organisation cannot reliably correlate demand, tool activity, and runtime consumption. That creates blind spots in cost attribution, threshold enforcement, and investigation.
Impact: Teams misread efficiency, miss emerging spend blowouts, and struggle to prove whether the agent was simply busy or genuinely misbehaving, which delays both financial control and operational response.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | GV.SC-01 — Cybersecurity Supply Chain Risk Management Strategy | Agent tool and worker usage create downstream operational and vendor-cost exposure that needs governance. |
| ID.AM-03 — Inventories of hardware managed by the organization are maintained | Layered agent usage tracking depends on maintaining accurate inventories of execution points and workers. | |
| PR.AA-05 — Access permissions and authorizations are managed, incorporating the principles of least privilege and separation of duties | Agent actions across tools and workers should be bounded and reviewed against authorised use. | |
| Recommendation — Define a usage governance strategy for agent tool and worker consumption. Maintain an inventory of agent, tool, and worker execution points. Enforce per-action authorization for agent tool and worker access. | ||
| NIST SP 800-53 Rev 5 | AU-6 — Audit Record Review, Analysis, and Reporting | Cross-layer usage tracking requires correlated audit data for spend and governance review. |
| AC-6 — Least Privilege | Agents that touch multiple tools and workers should only use the minimum access needed. | |
| Recommendation — Correlate agent, tool, and worker audit records for review. Limit each agent path to the minimum access required. | ||
Practitioner Guidance
What to prioritise: Build one reporting model that preserves the three layers separately before you try to optimise the agent. If you only tune the worker or only tune the tool layer, you can improve one metric while hiding the real cost driver.
What to verify: Make sure every agent run can be traced from the initiating user challenge to the tool calls and then to the worker runtime that executed them. If any hop is missing, the dashboard is not decision-grade.
What good looks like: Finance can explain spend movements, platform can explain capacity pressure, and security can explain unusual execution paths from the same telemetry set without rebuilding the story by hand.
Practitioner takeaway: The goal is not just to count agent activity, it is to preserve the boundary between demand, action, and execution so that cost, governance, and investigation stay trustworthy.
Related resources from NHI Mgmt Group
- What do teams get wrong about auditing AI agent activity across tools and workflows?
- What do teams get wrong about adopting AI tools across security, legal, finance, and engineering workflows?
- What do teams get wrong about AI agent access in MCP environments?
- What do security teams get wrong about AI agent identity governance?
Deepen Your Knowledge
Reviewed and updated by the NHIMG editorial team on September 30, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org