TL;DR: AI spend governance fails when invoices, provider consoles, and cost centres do not share a common identity and policy model, leaving teams unable to see who is spending, on what, and why, according to Stacklok. The real control problem is not budget capping alone, but governing access, attribution, and fallback without interrupting legitimate work.
At a glance
What this is: This is Stacklok’s analysis of what good AI spend governance looks like, and its key finding is that control only works when identity, usage, and policy are unified across providers, tools, and agents.
Why it matters: It matters because IAM, NHI, and governance teams need spend controls that attribute usage correctly, limit unmanaged credentials, and keep AI workloads operational without creating bypass paths.
By the numbers:
- The average enterprise now has 69% more machine identities than human ones.
- 59% of companies face greater difficulties auditing machine identities, primarily due to lack of clear ownership and limited visibility.
- Only 38% have automated certificate lifecycle management in place.
👉 Read Stacklok's analysis of AI spend governance and identity control
Context
AI spend governance is the discipline of making AI usage visible, attributable, and controllable across models, tools, and workloads. In practice, the problem is not just cost control. It is that finance, security, and engineering often see different slices of the same non-human identity estate, which makes it hard to know what is really being consumed or by whom.
That gap becomes more serious as organisations spread usage across direct provider accounts, cloud-hosted model services, developer tools, and agentic workflows. A governance model that cannot connect consumption to identity and organisational ownership will always produce partial answers, even when the raw invoices are accurate.
Key questions
Q: How should teams govern AI consumption when spend is spread across multiple tools?
A: Start by assigning one control owner for AI consumption governance and require shared evidence from finance, IT, and security. Then map usage back to the identity or workflow that generated it so spend, access, and accountability stay linked. Without that join, AI usage becomes visible only as a cost, not as governed behaviour.
Q: Why do AI budgets fail when they are based only on invoices?
A: Invoices report what was charged, not which identity consumed the service, which workload generated it, or whether the activity was intended. That means they are useful for accounting but weak for governance. A workable AI budget needs identity context, usage telemetry, and organisational ownership so the numbers can support decisions about reduction, reallocation, and control.
Q: What do organisations get wrong about AI spend visibility?
A: They often confuse partial dashboard coverage with complete governance. A view of one provider or one team can look authoritative while missing direct API keys, unmanaged agents, or cloud workloads that still spend money outside the control point. The right test is whether the control sits in front of the actual estate, not just the easiest part of it.
Q: How do you know if AI fallback policies are working?
A: They are working when substitutions are deliberate, visible, and limited to approved workloads. You should be able to see which model handled the request, why the fallback triggered, and whether the change preserved acceptable quality. If users cannot tell when a model switch happened, the policy is hiding cost decisions rather than governing them.
Technical breakdown
Why AI spend governance breaks across provider silos
AI spending is fragmented by design because providers expose different billing models, identity structures, and usage views. Direct model APIs, cloud-hosted services such as Bedrock, and enterprise offerings such as Azure OpenAI all report consumption differently, which makes a single financial view hard to construct. The problem gets worse when developer tools and agents use separate credentials or when usage bypasses the primary control point entirely. Governance therefore depends on inserting identity-aware policy in front of the estate, not just aggregating invoices after the fact.
Practical implication: establish a governed endpoint that normalises identity, policy, and usage across all model providers before you attempt chargeback or budget enforcement.
How hierarchical budgets support AI cost attribution
Hierarchical budgets are a governance model for allocating spend across individuals, teams, and shared pools. They work because they separate the question of who is spending from where the capacity is funded. That allows a personal allowance to be exhausted without automatically stopping the work if team or organisational capacity remains. The important technical detail is that budgets must be tied to identities, groups, and workload context, otherwise the accounting model cannot distinguish legitimate usage from waste or routing around policy.
Practical implication: design budget structures around identity and group ownership so you can attribute spend to the right organisational unit without forcing every request into a hard stop.
How model fallback changes policy from denial to controlled substitution
Policy-driven model fallback routes eligible requests from a higher-cost model to a lower-cost one when thresholds are reached. This is not the same as a blind failover, because the workload eligibility, destination model, and trigger condition must all be defined in advance. The architectural value is that governance can preserve service continuity while reducing cost, but only if the system makes the substitution visible and auditable. Otherwise the organisation cannot tell whether a response came from the intended model or a cheaper fallback.
Practical implication: define explicit fallback rules for low-risk workloads and require auditability of every model substitution so users and reviewers can see what changed.
NHI Mgmt Group analysis
AI spend governance is really NHI governance with a finance surface. The article shows that the control problem is not just invoice management, but attribution across models, tools, and workloads that are already acting as non-human identities. When credentials, usage, and organisational ownership are fragmented, finance sees cost while security sees only partial access. Practitioners should treat spend governance as an identity governance problem first and a cost problem second.
Partial visibility is more dangerous than obvious incompleteness. A dashboard that covers only some providers can create false confidence because it looks more operationally mature than a spreadsheet while still missing direct API keys, unmanaged agents, or cloud-hosted workloads. That is a control-plane issue, not a reporting issue. The practitioner lesson is to govern the full estate you actually have, not the subset that is easiest to integrate.
Identity-linked cost allocation is becoming a control requirement, not a finance convenience. Once spend can be mapped to users, groups, and workloads, security and compliance teams gain evidence that controls actually constrained usage rather than just moved it around. This is where NIST Cybersecurity Framework 2.0 and NHI governance intersect: visibility, accountability, and auditability have to survive cross-provider sprawl. The implication is that organisations need a defensible attribution model before they can trust any AI budget programme.
Controlled fallback is the right governance pattern when AI work must continue under budget pressure. Hard caps alone create bypass incentives, support churn, and hidden shadow usage. The stronger pattern is policy-driven substitution, where approved workloads can move to a lower-cost model without losing visibility or breaking accountability. That keeps the governance model aligned with how real teams work while preserving evidence for review.
From our research:
- The average enterprise now has 69% more machine identities than human ones, according to The Critical Gaps in Machine Identity Management report.
- Machine identity auditing remains weak, with 59% of companies reporting greater difficulty auditing machine identities because ownership and visibility are unclear.
- For the control side, Ultimate Guide to NHIs , Lifecycle Processes for Managing NHIs explains how lifecycle governance reduces unmanaged credential sprawl.
What this signals
AI spend governance will increasingly be judged on attribution quality, not just cost reduction. Teams that can map usage to identities, groups, and workloads will be able to distinguish legitimate growth from credential sprawl and route-around behaviour. That makes attribution the leading indicator for both financial discipline and security control.
Identity-linked budgets should become part of the same operating model as access reviews and offboarding. When people, workloads, and agents leave or change roles, their spending paths need to close with the same discipline as their access paths. Otherwise AI usage becomes another form of stale entitlement.
The next maturity step is to connect spend policy to workload identity governance and audit evidence. That means operational telemetry, budget thresholds, and approval paths need to converge inside a control plane that finance and security can both trust.
For practitioners
- Unify AI access through a governed endpoint Place model calls, tool calls, and agent traffic behind a single identity-aware control point so usage, policy, and audit context are consistent across providers.
- Map spend to organisational ownership before setting hard limits Tie budgets to users, groups, projects, or cost centres so finance can allocate cost and security can identify unmanaged credentials or workloads.
- Treat unmanaged API keys as spend and security risk Review developer laptops, application configuration, and direct provider access for credentials that can keep consuming after the creator has moved on.
- Define explicit model fallback rules for low-risk workloads Approve which workloads can shift to lower-cost models, which destination models are allowed, and when the switch must be visible to users and operators.
- Export identity-linked audit events to existing monitoring Send request identity, provider, model, timestamp, credential context, and cost into the SIEM or observability stack so reviews do not depend on manual invoice reconstruction.
Key takeaways
- AI spend governance fails when organisations treat invoices as the control rather than the output of a control model.
- The hard part is identity-linked attribution across providers, tools, agents, and workloads, not simply setting a lower ceiling.
- The strongest governance pattern is a single control plane that can see, budget, audit, and gracefully substitute without breaking legitimate work.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST CSF 2.0, NIST Zero Trust (SP 800-207) and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-03 | The article centres on unmanaged non-human credentials and governance gaps across AI workloads. |
| NIST CSF 2.0 | PR.AC-1 | Identity-aware access control is the basis for governing model, tool, and agent usage. |
| NIST Zero Trust (SP 800-207) | The governed endpoint model aligns with zero trust principles for every model and tool call. | |
| NIST SP 800-53 Rev 5 | AC-6 | Least privilege is needed when AI tools and workloads use direct API access or shared budgets. |
Bind AI usage controls to identity and access management under PR.AC-1 before allocating budgets.
Key terms
- AI Spend Governance: AI spend governance is the set of identity, policy, and accounting controls used to make AI consumption visible and accountable. It ties usage to people, teams, workloads, and cost centres so finance can allocate costs and security can prevent unmanaged access from becoming unmanaged spend.
- Identity Attribution: Identity attribution is the ability to determine which entity performed an action and under what authority. For AI agents, it requires separate identities, structured logs, and traceable decision records so investigations can distinguish human intent from autonomous execution.
- Hierarchical Budgeting: Hierarchical budgeting allocates AI capacity across multiple layers, such as individuals, teams, and shared organisational pools. It reduces unnecessary work stoppages because a local limit can be reached without ending all activity, provided higher-level capacity remains available and policy allows the request to continue.
- Model Fallback: Model fallback is a policy rule that routes eligible requests from one model to another when cost, capacity, or policy thresholds are reached. It is only governed well when the substitution is explicit, approved in advance, and visible to operators and reviewers.
What's in the full article
Stacklok's full blog insight covers the operational detail this post intentionally leaves for the source:
- How the AI Gateway normalises multiple provider APIs behind a common endpoint for identity and policy enforcement
- How token usage, gateway health, and rate-limiting activity flow into OpenTelemetry, Prometheus, and Grafana
- How model fallback is configured so approved workloads can shift to lower-cost models without losing auditability
- How local credential bridges reduce unmanaged keys when tools only support static API access
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
Published by the NHIMG editorial team on August 18, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org