By NHI Mgmt Group Editorial TeamDomain: AI SecuritySource: CycodePublished August 24, 2026

TL;DR: The AI security maturity model problem is not the framework itself but the reliance on self-report, because visibility is the first thing being scored and survey data shows 90% of organizations claim AI visibility while 59% still confirm or suspect shadow AI, according to Cycode. The practical fix is to bind every maturity claim to a retrievable artifact, then score only after discovery, baselining, and enforcement.


At a glance

What this is: This is an analysis of why AI security maturity models often mismeasure real risk, with the central finding that self-scored assessments fail when visibility is incomplete.

Why it matters: It matters because IAM, NHI, and AI governance teams need evidence-based controls for AI systems, agents, and tool access, not maturity grades built on unverified claims.

By the numbers:

👉 Read Cycode's analysis of why AI security maturity models need evidence, not self-report


Context

AI security maturity models are only useful when they measure evidence rather than optimism, because visibility gaps distort almost every self-scored assessment. In AI environments, the first question is not whether a team has policies, but whether it can produce artifacts that prove those policies are active across assistants, models, agents, and tool calls.

That is where the identity angle becomes real for IAM and NHI teams. AI systems increasingly depend on non-human identities, tool permissions, secret material, and delegated access paths, so maturity scoring that ignores identity governance misses the control plane that actually determines exposure.


Key questions

Q: How should security teams govern AI adoption when maturity scores look better than reality?

A: Security teams should anchor AI governance in identity and access controls, not self-assessed maturity. The practical test is whether the organisation can inventory AI-connected access, enforce approval, and revoke it cleanly. If those controls are fragmented, the maturity score is not a reliable indicator of safe scale.

Q: Why do AI maturity models fail when they rely on self-assessment?

A: Self-assessment fails because the first thing being measured is visibility, and AI environments are often only partially visible. Teams can claim governance, policy coverage, or inventory completeness without being able to prove it. Maturity scoring should therefore be treated as a validation exercise, not a trust exercise, especially when AI systems use secrets or non-human identities.

Q: What do organisations get wrong about AI-generated code as a control signal?

A: They often assume that clean-looking code or a passed review means the code is safe. In practice, AI-generated code can be persuasive, syntactically correct, and still introduce vulnerable dependencies or hidden supply chain risk. The better signal is whether the code is attributable, checked against your own baseline, and tied to measurable outcomes.

Q: Who should own AI governance when AI touches identity and access?

A: Ownership should sit with the team that can explain the AI system’s access, purpose, and operating boundaries end to end. In practice, that means AI governance must connect security, IAM, data, and engineering accountability so the system is not treated as a floating experiment. If ownership is unclear, lifecycle control will be inconsistent.


Technical breakdown

Why self-scoring fails when visibility is the control being measured

Most maturity models assume an organisation can accurately describe its own state before the assessment starts. In AI security, that assumption breaks because discovery is incomplete, inventories age quickly, and teams often score from memory rather than from system evidence. That means the maturity number reflects confidence, not control. This is especially dangerous where AI systems interact with secrets, tool permissions, or workload identities, because undiscovered access is indistinguishable from managed access if no artifact exists.

Practical implication: Treat discovery and baselining as prerequisites to scoring, not as follow-up tasks.

Why code review is a weak maturity signal for AI-generated code

AI-generated code can look clean, compile correctly, and still introduce risk that human reviewers miss. A maturity model that uses review coverage as proof of control is measuring an activity that no longer reliably detects defects or unsafe dependencies. The real issue is not whether code was reviewed, but whether the review process is tied to exploitability, attribution, and repository-level evidence. In agentic development, the same logic applies to NHI access paths that let code or tools act without sufficient scoping.

Practical implication: Anchor maturity claims to repository evidence, attribution, and defect outcomes rather than review volume.

Why AI maturity models need identity-aware evidence, not just policy statements

AI governance and identity governance are converging because agents, pipelines, and coding assistants depend on machine credentials, delegated permissions, and secret material. A model that ignores those identities will overstate maturity whenever policy exists on paper but access remains broad in practice. The useful shift is to score whether every AI-related access path has an owner, a retrievable authorization state, and a revocation trail. That is the difference between policy intent and operational control.

Practical implication: Require named ownership and retrievable access evidence for every AI-connected identity and tool.


Threat narrative

Attacker objective: The objective is to exploit weak visibility and over-trusted AI workflows to introduce risky code, dependencies, or access paths that bypass governance.

  1. Entry occurs through AI-assisted development, where assistants, MCP-connected tools, or generated dependencies enter the software factory without full inventory coverage.
  2. Escalation happens when those tools or credentials inherit broader access than intended, allowing code generation, dependency insertion, or tool actions to exceed review assumptions.
  3. Impact emerges as insecure code, hidden shadow AI, or ungoverned agent behaviour reaches production and creates security failures that the maturity score did not reveal.

NHI Mgmt Group analysis

AI maturity models are increasingly measuring governance theatre, not security maturity. When a scoring model relies on self-report, it rewards confidence and documentation quality more than actual control strength. That is tolerable in stable environments, but AI changes quickly and visibility is often partial, so the score becomes a summary of optimism. Practitioners should treat any maturity model as a hypothesis until it is backed by artifacts.

Shadow AI creates an evidence problem that maturity rubrics still underweight. If undiscovered assistants, agents, models, or MCP servers exist, then the organisation is scoring an incomplete inventory. That matters for IAM and NHI because undiscovered non-human access cannot be governed, reviewed, or revoked. The category needs discovery-first maturity thresholds, not policy-first ones.

Agentic development should be treated as an identity governance domain, not only a software delivery problem. AI systems that write code also consume secrets, call tools, and act through machine identities, so the control question is who or what is authorised to do those things. A maturity model that stops at AI policy misses the delegated access layer where real risk accumulates. Teams should map AI governance to NHI governance explicitly.

Evidence-based maturity models will outperform level-based scorecards because they force operational proof. A named person should be able to produce each artifact within one business day, and the score should remain blank when that proof does not exist. That approach is harder to game, easier to audit, and much closer to how boards and regulators assess control reality. Practitioners should build toward proof, not just progression.

AI governance debt is now a useful concept for this market. It describes the gap between the speed of AI adoption and the slower build-out of discovery, attribution, and enforcement evidence. The debt compounds each time an organisation adds tools, models, or agents without a corresponding control artifact. Practitioners should measure that debt before it becomes an incident response problem.

What this signals

AI security maturity will increasingly be judged on evidence quality, not framework selection. Teams that can produce inventories, attribution baselines, and blocked-action logs will have a defensible story; teams that cannot will keep overestimating progress. This is where identity control becomes central, because AI systems that act through machine credentials can only be governed when access state is visible and assigned.

The practical consequence for programmes is that AI governance, NHI governance, and secure development can no longer live in separate reporting lanes. If assistants, agents, and tool chains are writing code or taking actions, the organisation needs a single view of identity, permission, and enforcement across those paths. That is why the control discussion is moving toward proof, not posture.

Identity governance debt: the gap between AI adoption speed and the organisation's ability to prove who or what is authorised to act. This gap widens quickly when agents and tools are added faster than inventories, ownership, and revocation trails. Teams should measure that debt directly and use it to prioritise where to tighten control first.


For practitioners

  • Bind each maturity claim to an artifact Require a retrievable artifact for every scored statement, such as an AIBOM export, agent attribution report, or blocked-policy log. If nobody can produce the evidence within one business day, leave the field blank.
  • Score only after discovery and baselining Run a 90-day sequence that starts with discovery of assistants, models, MCP servers, packages, and AI secrets, then establishes your own repository baseline, and only then applies a maturity score.
  • Treat AI-connected identities as governed assets Inventory non-human identities, delegated tool permissions, and secret-bearing workflows alongside AI systems so that access ownership, authorisation state, and revocation paths are explicit.
  • Replace review volume with outcome evidence Compare AI-generated and human-authored code by attribution rate, violation rate, and production failure patterns instead of assuming a high review count equals strong control.
  • Map AI governance to NHI controls Assign owners to every AI tool path that can act, rotate, or request access, then align those paths with machine identity and least-privilege controls already used elsewhere in the programme.

Key takeaways

  • AI maturity models break down when they score self-report instead of evidence, because visibility is the control boundary that fails first.
  • The real governance gap is not the lack of a framework, but the lack of retrievable artifacts for AI inventories, attribution, and access state.
  • Identity-aware AI governance requires the same discipline used for NHI and privileged access: named ownership, least privilege, and provable revocation.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST AI RMFGOVERNThe article focuses on governance evidence and accountability for AI maturity scoring.
OWASP Agentic AI Top 10A1Agentic AI controls matter because the article covers agents, tools, and delegated action paths.
OWASP Non-Human Identity Top 10NHI-03Non-human identities and secret-bearing workflows are central to AI systems that act autonomously.
NIST CSF 2.0PR.AC-1Access control and identity governance underpin the evidence-based maturity approach described here.
NIST SP 800-53 Rev 5IA-5Authenticator management is directly relevant where AI systems rely on secrets and machine credentials.

Tie AI maturity claims to access inventories and authorisation evidence before scoring programme posture.


Key terms

  • AI maturity model: A structured framework for assessing how far an organisation has progressed in adopting and operationalising AI. In practice, it helps teams compare pilots, production use, and enterprise-wide integration by looking at governance, data quality, capability, and lifecycle discipline rather than adoption hype.
  • Shadow AI: AI agents, copilots, or connected tools operating without full visibility or governance from security teams. Shadow AI becomes an identity problem when those systems authenticate with unmanaged tokens, service accounts, or OAuth apps that can reach production resources.
  • AIBOM: An AI Bill of Materials is a structured inventory of the components, data sources, prompts, connectors, and dependencies that shape an AI system. It helps security teams understand what the model can access, where risk enters the stack, and which changes require governance review.
  • Agent Attribution: Agent attribution is the ability to tie each request, action, retry, and cost event back to a specific non-human actor. It matters because shared credentials and anonymous automation hide misuse, make revocation harder, and prevent security teams from understanding which identity actually performed a task.

What's in the full article

Cycode's full analysis covers the operational detail this post intentionally leaves for the source:

  • The comparison logic between the five AI security maturity models and the specific questions each one is best suited to answer.
  • The artifact-by-artifact scoring method, including the evidence expected for inventories, attribution baselines, and enforcement claims.
  • The 90-day discovery, baselining, and enforcement sequence for turning a maturity rubric into something auditable.
  • The vendor's own research context, including codebase visibility data and the implications for product security teams.

👉 Cycode's full post covers the maturity model comparisons, evidence rubric, and 90-day operational sequence.

Deepen your knowledge

The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It helps practitioners connect identity controls to the broader access and lifecycle questions that AI governance now raises.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 25, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org