TL;DR: Data visibility can show where sensitive information lives, but it does not reduce exposure if users, applications, service accounts, and AI systems can still reach it, according to BigID. The real security gap is access governance, because controlling who can use data now matters more than simply finding it.
At a glance
What this is: This is an analysis of why data discovery and classification do not equal security when access remains broad, persistent, and hard to govern.
Why it matters: It matters to IAM practitioners because data protection now depends on identity context, entitlement scope, and usage control across human and non-human access paths.
By the numbers:
- 72% of organisations have experienced or suspect they have experienced a breach of non-human identities, 46% confirmed and 26% suspected.
👉 Read BigID's analysis of why data access, not data location, defines risk
Context
Data visibility tells teams where information exists, but access governance determines who can use it and under what conditions. In modern environments, that distinction matters because sensitive data is rarely exposed by location alone. It is exposed when privileges expand beyond intent across users, roles, service accounts, applications, and AI systems.
This is where identity becomes part of data security. If organisations can classify data but cannot map effective access, they miss the control layer that actually reduces risk. That makes the article relevant to IAM, PAM, NHI governance, and AI access oversight because access paths now span both human and machine identities.
Key questions
Q: How should security teams prevent unauthorized access across human and machine identities?
A: They should use different controls for interactive users and non-human identities, but govern them through one access model. Enforce MFA and strong recovery for people, then inventory API keys, tokens, service accounts, and third-party access separately. The key is to remove standing privilege, shorten credential lifetime, and review delegated access on a fixed schedule.
Q: Why do organisations struggle to secure data even when classification is mature?
A: Classification tells you what data is sensitive, but not whether the right identities can reach it. In complex environments, permissions drift across SaaS, cloud, and machine identities faster than teams can review them. The result is a gap between knowing the value of the data and controlling the pathways that expose it.
Q: What do security teams get wrong about data volume and visibility?
A: Teams often assume that more telemetry automatically means better visibility, but data volume without purpose just increases noise. Some sources are valuable for retention or audit, while others are meant to drive detection and response. The mistake is treating every event stream as if it should produce alerts.
Q: How can organisations tell whether governed data access is actually working?
A: Look for fewer shadow copies, faster request fulfilment, consistent metric definitions and lower variation in how teams consume the same data. If users still create duplicate sources of truth, the governance model is not enabling trusted access. Effective control shows up in reduced friction and higher confidence, not just more policy documentation.
Technical breakdown
Why data discovery does not equal data control
Discovery and classification tools answer a storage question: where is sensitive data located, and how sensitive is it? They do not answer the governance question that matters operationally: which identities can read, transform, copy, or expose that data. In practice, the same dataset may be reachable through users, groups, applications, service accounts, or AI workflows. Without correlating entitlement state to actual usage, security teams create visibility without enforcement, which leaves overexposure hidden even in well-instrumented environments.
Practical implication: tie discovery outputs to entitlement and activity data before treating a dataset as secured.
Data access governance in identity-rich environments
Data access governance extends beyond RBAC by adding identity context, usage patterns, and policy intent to access decisions. RBAC can tell you who should have access in theory, but it often fails when roles accumulate, applications inherit broad permissions, or machine identities are created faster than reviews occur. The result is access drift. That drift is especially dangerous in environments with SaaS integrations, cloud storage, and AI systems because each new access path can widen the blast radius without changing the underlying data classification.
Practical implication: review effective access, not just assigned roles, for users, service accounts, and applications.
Why AI systems turn data access into a governance issue
AI systems intensify the access problem because they do not merely store or display data, they query it, transform it, and surface it through outputs and downstream workflows. If an AI system can reach sensitive sources without tightly scoped permissions, it can expose data that would otherwise remain compartmentalised. This makes AI access a non-human identity problem as much as a data protection problem. The governance challenge is not the model itself alone, but the data pathways and permissions that the model can invoke.
Practical implication: govern AI data reach with least privilege, scoped tokens, and monitored access paths.
Threat narrative
Attacker objective: The objective is to reach sensitive data through legitimate access paths and expand that access into broader exposure or misuse.
- Entry begins when over-permissioned users, applications, or machine identities gain access to sensitive datasets through broad entitlements or inherited permissions.
- Escalation occurs when those identities retain access longer than necessary, allowing copying, lateral reuse, or indirect exposure across systems and workflows.
- Impact follows when sensitive data is misused, overexposed, or surfaced through AI outputs, downstream integrations, or uncontrolled sharing.
NHI Mgmt Group analysis
Data visibility without access governance creates a control illusion. Organisations can know exactly where sensitive data lives and still fail to reduce risk if entitlement scope is unmanaged. The issue is not discovery quality, but the absence of effective control over who can use data, especially across service accounts and applications. Practitioners should treat visibility as input to enforcement, not evidence of protection.
Access is now the primary risk layer in data security. As environments grow more identity-rich, permissions accumulate faster than teams revoke them. That means risk is driven less by the presence of sensitive data and more by the persistence of unnecessary access. This aligns with NIST CSF 2.0 and NIST SP 800-53 controls around access governance and auditability, and practitioners should measure access drift as a core security signal.
Machine identities are central to data exposure, not peripheral to it. The article correctly points to applications and AI systems because many data access paths now run through non-human identities. That makes NHI governance part of data protection architecture, not a separate discipline. The named concept here is access-driven exposure: sensitive data becomes risky when identity permissions outgrow business necessity. Teams should govern that exposure before classification efforts become cosmetic.
AI security begins with data reach, not model behaviour. If an AI system can query or transform sensitive data without strict authorisation boundaries, the model becomes a propagation layer for existing access weaknesses. The governance question is therefore who or what can invoke data, under what context, and with which revocation rules. Practitioners should align AI access with identity lifecycle controls and least-privilege enforcement.
Data access governance is becoming the missing bridge between DSPM and IAM. DSPM can identify sensitive data, and IAM can define identities, but neither alone explains effective access across real workflows. The stronger programme combines identity context, policy intent, and observed usage into one control view. Practitioners should expect data security programmes to converge with IAM, PAM, and NHI controls rather than remain separate.
What this signals
Access-driven exposure is the practical problem this topic exposes, because data programmes still over-rely on discovery while under-investing in entitlement and usage governance. For readers running identity-heavy environments, the next step is to join DSPM outputs with IAM, PAM, and NHI controls so sensitive data cannot remain reachable by stale permissions or unmonitored machine identities. The NIST Cybersecurity Framework 2.0 remains the cleanest operational frame for that shift, especially where governance and access control need to be measured together.
AI systems make this problem sharper, not softer, because data access now flows through software entities that can act at speed and at scale. Where teams already use the [NHI Lifecycle Management Guide](https://nhimg.org/nhi-lifecycle-management-guide) to manage provisioning and offboarding, the same logic should extend to AI data reach and service account scope. Access that cannot be explained, reviewed, and removed is already outside governance boundaries.
Usage-aware governance: the next maturity step is correlating identity, permissions, and activity in one control view. That gives security and compliance teams evidence for least privilege, data minimisation, and revocation decisions instead of relying on data labels alone. When access signals and data sensitivity live in separate tools, the organisation will keep discovering risk faster than it can reduce it.
For practitioners
- Map effective access to sensitive datasets Correlate data discovery results with assigned and observed access across users, roles, applications, service accounts, and AI systems so teams can see who can actually use the data.
- Review entitlement drift on a fixed cadence Identify roles, groups, and machine identities whose permissions exceed current business need, then remove access that has persisted beyond its operating purpose.
- Treat AI data access as non-human identity governance Require scoped credentials, explicit purpose boundaries, and monitored access paths for AI systems that query or transform sensitive information.
- Use usage signals to validate least privilege Track whether access aligns with actual data use, then investigate identities that access sensitive records without a clear operational need or expected workflow.
Key takeaways
- Data discovery is necessary, but it does not secure sensitive information when access remains broad.
- The biggest governance gap is not knowing where data lives, but knowing which identities can actually use it.
- Identity context, usage telemetry, and least-privilege enforcement are the controls that turn visibility into real risk reduction.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST CSF 2.0, NIST SP 800-53 Rev 5, NIST Zero Trust (SP 800-207) and NIST AI RMF set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | PR.AC-4 | Access management is central to the article's argument about controlling who can use data. |
| NIST SP 800-53 Rev 5 | AC-6 | Least privilege directly addresses the overexposure problem described in the article. |
| OWASP Non-Human Identity Top 10 | NHI-03 | Unmanaged and excessive machine access is part of the data exposure problem. |
| NIST Zero Trust (SP 800-207) | Section 2.1 | Zero Trust supports continuous verification of access to sensitive data across identities. |
| NIST AI RMF | GOVERN | AI data access governance needs clear accountability and policy ownership. |
Use NHI-03 to locate service accounts and AI identities that can reach sensitive datasets without tight scope.
Key terms
- Data Access Governance: Data access governance is the practice of deciding who or what should reach specific data based on sensitivity, business purpose, and observed access paths. It combines classification, entitlement analysis, and review workflows so access decisions reflect exposure, not just permission status.
- Permission Drift: Permission drift is the gradual expansion of access beyond what was originally intended. It happens when roles, tokens, and service accounts accumulate unused rights over time, making cloud identities harder to review and more dangerous to compromise.
- Non-Human Identity (NHI): A digital identity assigned to a non-human entity such as a software application, service account, API key, bot, machine, or AI agent that enables it to authenticate and interact with systems without direct human involvement. NHIs now outnumber human identities in most enterprises by 25 to 50 times.
- Access-Driven Exposure: Access-driven exposure is the condition where sensitive data becomes risky because identities can reach it too broadly or too persistently. The issue is not the dataset itself, but the permissions, workflows, and machine connections that make it usable outside intended boundaries.
What's in the full article
BigID's full analysis covers the operational detail this post intentionally leaves for the source:
- How BigID correlates sensitive data discovery with user, role, and non-human identity access paths
- Examples of overexposed datasets and the access patterns that keep them reachable
- The mechanics of combining DSPM with access governance in one workflow
- BigID's data access risk framing for teams moving from visibility to control
👉 The full BigID article explains how data access governance and DSPM are connected in practice.
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, identity lifecycle, secrets management, and workload identity. It gives security and identity practitioners a practical foundation for governing both human and non-human access paths.
Published by the NHIMG editorial team on August 20, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org