TL;DR: AI token prices are falling, but usage is rising faster, so token spend no longer tells teams whether AI is producing value; Arize argues for measuring cost per outcome instead, using traced and evaluated agent runs to connect spend to resolved tickets, accepted PRs, or shipped features. The practical shift is from billing visibility to outcome governance, where cost efficiency is judged by work actually completed rather than tokens consumed.
At a glance
What this is: This analysis argues that token-based AI cost metrics are a poor proxy for value because usage can rise without a corresponding increase in shipped outcomes.
Why it matters: IAM, NHI, and AI governance teams need outcome-based measurement because autonomous or semi-autonomous systems can consume access, compute, and secrets without proving business value.
👉 Read Arize's analysis of why token costs do not show whether AI is working
Context
AI spend is increasingly easy to count and increasingly hard to interpret. Token bills show activity, but they do not show whether an AI system is resolving work, changing code, or improving service delivery. That creates a governance gap for teams responsible for AI controls, because cost growth alone does not prove effectiveness.
For identity and access programmes, the issue matters because AI systems often operate through credentials, tools, and delegated access. When those systems are treated as productive simply because they consume more tokens, organisations can miss the real control question: whether the identity, access, and runtime permissions granted to the system are producing measurable outcomes or just creating more exposure.
Key questions
Q: How should teams measure whether an AI workflow is actually working?
A: Measure AI cost against a verified outcome such as a resolved ticket, accepted pull request, or shipped feature. Token usage shows consumption, but only outcome metrics show whether the system produced durable value. Add traces and evaluation labels so the metric reflects completed work rather than activity alone.
Q: Why do token metrics fail as a governance signal for AI systems?
A: Token metrics fail because they measure volume, not usefulness. An AI system can consume more context, more calls, and more budget while still creating rework or no lasting output. Governance needs a metric that links delegated access and spend to measurable business results.
Q: What do security teams get wrong about AI cost control?
A: They often treat cost as a finance-only issue and overlook the identity layer that drives usage. Without attribution, shadow AI discovery, and policy enforcement at the request path, teams can reduce waste in one area while leaving the real source of token growth untouched.
Q: Who should own cost per outcome reporting for AI programmes?
A: Ownership should sit across finance, engineering, and security, because the metric joins spend, delivery, and control effectiveness. Security and identity teams should ensure the system of record includes traces and access context, while product or engineering teams validate the business result.
Technical breakdown
Why token cost is a weak control metric for AI workloads
Token cost measures consumption, not effect. A model or agent can burn through context, calls, and prompt tokens while producing nothing durable, or it can use fewer tokens and still create material value. That makes token spend a lagging billing metric rather than a governance signal. In operational terms, token price tells you what inference costs, but not whether the work completed, whether the output was accepted, or whether the system improved the process it was meant to support. For teams governing AI access, this is analogous to tracking login volume instead of privileged action success. The unit of cost must match the unit of value.
Practical implication: define the business outcome first, then measure AI cost against that outcome instead of against raw token consumption.
What cost per outcome measures in AI operations
Cost per outcome divides total AI spend by a verified result such as a resolved ticket, accepted pull request, passing test, or shipped feature. That makes the metric useful because it aligns finance, engineering, and governance around the same evidence of value. The hard part is that most internal workflows do not expose clean success signals by default, so teams need traces, evaluation labels, and outcome definitions before the ratio is trustworthy. Without that, a model can look efficient while still generating rework, human review, or silent failures. Outcome-based accounting turns AI from an opaque expense into a measurable production system.
Practical implication: instrument agent traces and label success criteria before using AI spend data in budget or control decisions.
Why evaluation layers are now part of AI governance
Evaluation is the mechanism that makes outcome-based measurement possible. An observability layer captures each agent run, while evals decide whether the output met the task objective. That pairing is essential for AI governance because access, autonomy, and spending all become harder to justify if success is never measured. For identity and NHI practitioners, the intersection is clear: AI systems use service accounts, API keys, and delegated permissions to act, so governance must cover both what they can access and what they actually accomplish. Otherwise, organisations end up expanding machine privileges without evidence that the machine is producing anything worth the risk.
Practical implication: tie every privileged AI workflow to an evaluation signal so access review and value review can happen together.
NHI Mgmt Group analysis
Token spend creates governance blindness when it is treated as a proxy for value. AI systems can consume large volumes of compute, prompts, and delegated access while producing no measurable business outcome. That is not just a finance problem, because access and runtime permissions are being exercised inside systems whose usefulness has not been proven. The result is a control gap where activity is mistaken for effectiveness, and effectiveness is what justifies privilege. Practitioners should treat token cost as an accounting input, not an assurance signal.
Cost per outcome is a more defensible control concept because it aligns spend with verified work. A resolved ticket, accepted PR, or shipped feature is a governance event, not just a usage event. Once AI systems start acting through service accounts and tool access, the programme needs a metric that can stand up in both security review and budget review. That is why outcome measurement belongs inside AI governance, not outside it. Practitioners should define outcome thresholds before expanding machine access.
Outcome-based measurement exposes a named concept we can call value-to-usage drift. This is the gap between rising AI consumption and flat or unproven business output. It matters because cost optimisation, privilege scope, and model selection all become distorted when usage rises faster than validated results. In an identity context, the same drift can hide over-privileged AI workflows that are busy but not useful. Practitioners should require evidence of outcome lift before increasing access, spend, or autonomy.
AI governance now depends on proving that delegated access produces durable results. When an agent uses credentials to act, the real question is whether that action changed something the business values. That puts evaluation, traceability, and access control in the same control plane. For teams managing IAM, PAM, and NHI programmes, the implication is that machine identity governance cannot stop at authentication or secrets management. Practitioners should connect permissioning to measured outcomes, or they will keep funding invisible waste.
NIST AI RMF and identity governance are converging at the point of measurable accountability. The AI RMF asks organisations to manage risk across the AI lifecycle, while identity programmes control what systems can access and do. Cost per outcome is the bridge between them because it shows whether the granted access generated acceptable value. Practitioners should use that bridge to decide when a model, agent, or workflow deserves broader permissions versus tighter constraints.
What this signals
Cost per outcome will become a governance requirement, not a performance preference. As AI systems take on more delegated work, programme owners will need evidence that the access they grant produces measurable value. That shifts the discussion from cheaper tokens to defensible control over machine activity, evaluation, and privilege.
For identity-led teams, the next control boundary is not just who or what can authenticate, but whether authenticated machine activity earns its access. That means IAM, PAM, and NHI governance need shared reporting with evaluation and observability teams, especially where agents can call tools, open tickets, or change code.
Value-to-usage drift: this is the condition where AI consumption rises faster than validated business output, masking waste and overreach. Teams should watch for it whenever usage increases without a corresponding change in delivery metrics, because that pattern often signals that access, autonomy, or model selection needs rework.
For practitioners
- Define outcome metrics before expanding AI usage Choose one or two business outcomes per workflow, such as resolved tickets, accepted PRs, or shipped features, and make them the denominator for spend analysis.
- Instrument every agent run with traceability Capture prompts, tool calls, and permissioned actions so each run can be linked to a specific task and reviewed against its result.
- Separate value review from usage review Review token growth, outcome success, and access scope as different controls so a high-usage system is not automatically treated as a successful one.
- Tie privileged AI access to measured performance Require evidence that a workflow improves the target outcome before granting broader permissions, longer retention, or higher autonomy.
Key takeaways
- Token spend alone cannot tell organisations whether AI is creating value or just consuming budget.
- Outcome-based measurement gives security, identity, and finance teams a common way to judge whether AI access is justified.
- Traceability and evaluation are now part of AI governance because delegated machine activity must be tied to verified results.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI RMF, NIST SP 800-53 Rev 5 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | MEASURE | The article centres on measuring AI effectiveness and risk, which maps directly to the Measure function. |
| NIST SP 800-53 Rev 5 | AU-2 | Trace capture and auditability are necessary to attribute AI runs to outcomes. |
| OWASP Agentic AI Top 10 | Outcome-based control depends on governing how agents use tools and complete tasks. | |
| NIST CSF 2.0 | GV.OV-01 | Governance requires oversight metrics that show whether AI investment is producing value. |
| MITRE ATLAS | TA0007 , Discovery | AI observability and trace review help identify how agent behaviour diverges from intended outcomes. |
Define oversight metrics that connect AI spend to business outcomes and control effectiveness.
Key terms
- Cost per outcome: A measure of the full cost required to complete a business task, including retries, review, exceptions, and rework. It is more useful than token cost alone because it reflects whether an AI workflow is actually efficient at scale.
- Outcome Metric: An outcome metric measures whether a security or identity programme changed the real-world state it was meant to influence. For NHI and IAM work, that means reduced exceptions, fewer repeated findings, faster remediation, or lower exposure, not just more completed tasks.
- Agent Trace: A structured record of an AI agent’s runtime activity, including model calls, tool calls, approvals, and subagent steps. In practice, traces support debugging, evaluation, and governance when they are retained, searchable, and tied to the permissions behind the agent.
- Value-to-Usage Drift: Value-to-usage drift is the gap that appears when AI consumption rises faster than verified business output. It is a governance problem because it can hide waste, overuse of access, and weak model selection behind apparently healthy activity.
What's in the full article
Arize's full article covers the operational detail this post intentionally leaves for the source:
- The cost-per-outcome implementation logic behind trace capture and evaluation scoring for AI workflows.
- Examples of how teams can map spend to resolved tickets, accepted pull requests, and shipped features.
- The observability workflow used to turn agent runs into auditable performance and budget data.
- The practical distinction between token-based billing, outcome-based pricing, and outcome-based governance.
Deepen your knowledge
NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It helps practitioners connect delegated access, runtime control, and lifecycle governance to broader security programmes.
Published by the NHIMG editorial team on August 19, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org