TL;DR: AI agent risk registers break when they are built around behaviour, because prompt changes, new MCP connections, and model upgrades make yesterday’s likelihood and impact ratings obsolete, according to ARMO. The durable unit is the agent surface, not the specific action, and runtime telemetry is now a prerequisite for usable governance.
At a glance
What this is: This is a risk-management framework for cloud-deployed AI agents that replaces behaviour-based registers with stable Identity, Data, Tool, and Model surfaces.
Why it matters: It matters because IAM, GRC, and security teams need a register structure that survives non-deterministic agent behaviour, new tool connections, and changing runtime scope.
By the numbers:
- 80.9% of teams have moved into active testing or production while only 14.4% have full security approval for their AI agent fleet.
- 88% have already had confirmed or suspected incidents.
👉 Read ARMO's framework for AI agent risk registers in the cloud
Context
AI agent risk management fails when security teams try to catalogue behaviour instead of the identity, data, tool, and model surfaces that shape it. In cloud-deployed AI agents, the execution path can change after a prompt shift, a new MCP connection, or a model update, while the underlying governance surfaces remain visible enough to manage.
That is the core problem for AI agent identity governance: conventional registers assume a stable threat, stable impact, and stable control set. For agents, those assumptions collapse quickly, which means the register has to be built around the parts of the system that remain stable enough to review, certify, and measure.
The article is typical of the current state of AI governance thinking: useful structure, but still early-stage operational maturity. The next step is not more labels; it is runtime evidence tied to identity scope, tool exposure, and downstream reach.
Key questions
Q: How should security teams build an AI agent risk register that survives changing behaviour?
A: Anchor the register to stable surfaces, not to the agent’s latest behaviour. Identity, Data, Tool, and Model are durable enough to review across prompt changes, tool additions, and model upgrades. That makes the register a governance artifact that can be maintained in production rather than rewritten every time execution shifts.
Q: Why do AI agents break traditional likelihood and impact scoring?
A: Because the underlying inputs are no longer stable enough for intuition-based scoring. Exposure, autonomy, instrumentation gap, identity blast radius, data sensitivity, tool capability, and model reach can be observed and measured. When those inputs are used consistently, different reviewers can reach the same risk tier.
Q: How do you know if an AI agent risk register is actually working?
A: A working register stays current after model upgrades, new MCP connections, and prompt changes. It also produces repeatable scores, mapped treatments, and monthly KRIs that match runtime evidence. If the rows only make sense on the day they were written, the register is not governing anything.
Q: Who should own AI agent governance when identity and access are shared across teams?
A: AI agent governance should sit with identity, security, and platform owners together, because no single team sees the full risk surface. IAM owns the control model, security owns containment and monitoring, and platform teams own the runtime integration. Shared ownership matters because agent risk spans identity, policy, and downstream execution.
Technical breakdown
Why behaviour-based AI agent risk rows fail
Traditional registers try to pin risk to a specific behaviour, but AI agents do not stay behaviourally fixed. A new prompt, a model upgrade, or an added MCP connection changes what the agent can do without changing the agent itself. That makes behaviour a poor row key for governance. The article’s central insight is that risk must be anchored to stable surfaces: Identity, Data, Tool, and Model. Those surfaces hold their shape even when execution varies. Practical oversight therefore depends on mapping the agent to its surfaces, not rewriting the register every time runtime behaviour shifts.
Practical implication: Use agent × surface rows as the register unit, not agent × observed behaviour.
How runtime-derived inventory changes the control model
AI agent inventory needs to be derived from runtime evidence because declared inventories miss what actually runs. A runtime AI bill of materials captures the agent process, the model endpoints it calls, the network connections it opens, and the API patterns it uses. That gives security teams a control plane for governance that is grounded in observable execution rather than CMDB completeness. It also separates the four surfaces cleanly: identity bindings, reachable data, exposed tools, and model/output channels. The value is not discovery alone. The value is making each surface independently reviewable, auditable, and updateable.
Practical implication: Build runtime discovery into the register pipeline so new tools and permissions are visible before they become exposure.
Why likelihood and impact need observable inputs
The article replaces subjective 1-to-5 scoring with inputs that can be measured. Likelihood becomes Exposure, Autonomy, and Instrumentation Gap. Impact becomes Identity Blast Radius, Data Sensitivity, Tool Capability, and Model Reach. This is a governance advance because it forces the committee to score the same properties every time rather than argue over intuition. It also aligns with NIST AI Risk Management Framework thinking, where the measurement problem matters as much as the control problem. For security leaders, the practical shift is from opinion-based risk rating to repeatable, evidence-backed classification.
Practical implication: Standardise scoring around observable deployment properties so different reviewers reach the same classification.
NHI Mgmt Group analysis
Behaviour is the wrong governance unit for AI agent risk. Traditional registers assume that a risk row can be tied to a stable action, but AI agents change execution paths as prompts, tools, and models change. That makes the action itself non-durable as a governance object. The useful unit is the surface that carries the behaviour, which is why Identity, Data, Tool, and Model are the right abstraction for IAM and GRC teams. The implication is that risk governance for agents has to move from event logging to surface control.
Runtime evidence is now part of identity governance, not just detection. When agent permissions, tool connectivity, and output paths are only known at design time, the programme is already behind. Runtime-derived inventory turns identity scope into something reviewable in production, which is where cloud-deployed agents actually operate. That shifts the identity programme from periodic approval to continuous scope verification. Security leaders should treat runtime telemetry as a governance input, not merely a security monitoring feed.
Likelihood and impact scoring become credible only when they are operationalised. The article’s Likelihood × Impact model matters because it replaces analyst intuition with observable inputs such as exposure, autonomy, instrumentation gap, blast radius, and data sensitivity. That is the kind of structure a risk committee can defend. It also creates a common language across IAM, cloud, and AI governance teams, which is essential when agents span all three. The practitioner conclusion is simple: if the inputs are not measurable, the risk score is not governable.
Identity blast radius is the named concept security teams should adopt. In agent programmes, the real question is not whether the model is accurate enough. It is how far the agent’s identity can reach, what data it can touch, what tools it can invoke, and which downstream systems trust its output. That blast radius is what drives escalation, and it is what IAM teams must constrain if the register is to mean anything operationally. The conclusion is that agent governance should be organised around blast radius, not around model novelty.
AI agent governance now sits at the intersection of IAM, cloud security, and AI risk management. The article correctly treats the register as a programme artifact, not a one-off checklist. That matters because agents can inherit privilege from service accounts, expand tool reach through MCP, and create new downstream trust chains through model outputs. No single control domain sees that whole picture. The conclusion is that cross-functional ownership is not optional if the register is to stay current.
From our research:
- 85% of organisations lack full visibility into third-party vendors connected via OAuth apps, according to The State of Non-Human Identity Security.
- Lack of credential rotation is cited as the top cause of NHI-related attacks by 45% of organisations, which shows how often governance failures start with identity lifecycle drift.
- For broader lifecycle framing, see Ultimate Guide to NHIs , Lifecycle Processes for Managing NHIs for how provisioning, rotation, and offboarding work together.
What this signals
Identity blast radius will become a core programme metric for AI agent governance. Once agents can add tools or expand output channels during runtime, the meaningful unit of control is no longer the model alone. Security teams should prepare to report on how far an agent identity can reach across data, tools, and downstream systems, not just whether it passed a pre-launch review.
The governance gap is structural: programme teams that rely on periodic review cycles will keep missing the moment when an agent’s scope changes. That is why runtime telemetry, surface-based registers, and control ownership across IAM and cloud teams matter more than a single approval checkpoint.
For practitioners aligning this with broader standards, the NIST AI Risk Management Framework gives the language for measurement and governance, while the OWASP Agentic AI Top 10 helps frame tool misuse and identity abuse in operational terms.
For practitioners
- Map every AI agent to four stable surfaces Rebuild the register around Identity, Data, Tool, and Model for each production agent, and update only the affected cells when model or tool changes occur.
- Replace subjective scoring with observable inputs Score likelihood from exposure, autonomy, and instrumentation gap, then score impact from identity blast radius, data sensitivity, tool capability, and model reach.
- Add runtime-derived inventory to GRC workflows Feed runtime AI bill of materials data into the register so the platform team can show actual agent processes, model endpoints, and API use rather than declared intent.
- Define KRI thresholds before deployment Set monthly thresholds for behavioural deviation, permission entropy, and cross-agent correlation before the agent reaches production, so escalation is pre-agreed.
Key takeaways
- AI agent risk management fails when security teams organise around behaviour instead of stable governance surfaces.
- Runtime evidence, not declared inventory, is what makes identity scope, tool exposure, and downstream reach reviewable.
- Programmes that cannot measure exposure, autonomy, and blast radius will not be able to govern AI agents at scale.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | The article focuses on agent behaviour, tool use, and runtime governance. | |
| OWASP Non-Human Identity Top 10 | NHI-03 | The article’s identity surface and permissions model are classic NHI governance issues. |
| NIST AI RMF | MEASURE | The scoring model depends on observable inputs and repeatable risk measurement. |
| NIST CSF 2.0 | PR.AC-4 | The framework centres on access scope and identity blast radius. |
| NIST Zero Trust (SP 800-207) | 3.1 | Runtime verification and scope control align with continuous trust validation. |
Map agent tool exposure and runtime behaviour to agentic AI risk categories before production approval.
Key terms
- AI Agent Risk Register: A structured record of AI agents, their exposures, controls, and residual risk. In practice, it should be built from stable governance surfaces such as identity, data, tool, and model, so that changing runtime behaviour does not force a full rewrite every time execution changes.
- Runtime-derived inventory: An asset inventory created from what is actually running and connecting in production, rather than what was declared at design time. For AI agents, it captures processes, model endpoints, API use, and network paths, which makes governance evidence-based instead of aspirational.
- Identity Blast Radius: The amount of damage a compromised identity can cause across systems, data, and infrastructure. In NHI environments, it is shaped by permissions, network reach, and administrative capability rather than by the credential alone. Reducing blast radius is a containment strategy that limits lateral movement and data exposure.
- Instrumentation gap: The difference between what an organisation can observe about an agent and what the agent can actually do. A large gap weakens likelihood scoring, hides scope drift, and makes risk committees depend on guesses rather than telemetry.
What's in the full article
ARMO's full blog covers the operational detail this post intentionally leaves for the source:
- A complete scoring model for Exposure, Autonomy, and Instrumentation Gap across cloud-deployed agents.
- The full treatment of surface-by-surface treatments for Identity, Data, Tool, and Model controls.
- Monthly KRI examples for behavioural deviation, permission entropy, and cross-agent correlation.
- Deployment guidance for turning runtime telemetry into a maintained AI agent register.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity security are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an identity security programme, it is worth exploring.
Published by the NHIMG editorial team on August 19, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org