TL;DR: AI applications are expanding enterprise data exposure by pulling in emails, chat logs, legal documents, and cloud files, while legacy classification and access controls struggle to keep up, according to Cyera. The governance problem is now less about storing data than knowing what AI can reach, classify, and expose before that reach becomes a breach.
At a glance
What this is: This is Cyera's analysis of why DSPM has become central to AI security, with data visibility, classification, and access governance now defining exposure risk.
Why it matters: IAM, NHI, and data security teams need the same visibility into AI tool access paths that they already expect for privileged humans and service identities.
By the numbers:
- The total volume of data is expected to reach 181 zettabytes in 2025.
- Cyera says its classification approach achieves 95% precision.
Context
AI security in this article is really a data governance problem. Cyera's point is that when generative AI can reach emails, chat logs, legal documents, file shares, and other unstructured content, security teams lose track of what the system can see before they can decide what it should see.
Legacy classification and access control models were built for slower-moving, more structured data estates. The article argues that DSPM closes that gap by discovering data across environments, classifying unstructured content, and showing where AI tools have reach that exceeds intended business access.
For IAM and NHI programmes, the practical issue is not only whether the AI tool is authenticated, but whether its access scope is understood, justified, and monitored against the sensitivity of the data it can ingest or expose.
Key questions
Q: How should security teams govern sensitive data used by AI systems?
A: Security teams should treat AI as a data consumer that needs policy boundaries, not just authentication. Classify sensitive data, define which datasets may enter AI workflows, and monitor outputs, logs, and downstream reuse. If governance stops at login, the organisation can approve access while still losing control of the data itself.
Q: Why do legacy data classification tools fail for generative AI use cases?
A: They were built for more structured data and rely heavily on manual tagging or pattern matching, which performs poorly on unstructured content. Generative AI consumes the messy material those methods often misclassify, so false positives and missed sensitivity decisions become a governance problem.
Q: What is the difference between access control and data governance in AI environments?
A: Access control decides which identities may reach a system or dataset. Data governance decides how data is classified, monitored, and approved for use. In AI environments, those two disciplines must work together because model input, output, and action paths can expose sensitive data even when the source dataset appears well governed.
Q: What should organisations prioritise first when securing AI data exposure?
A: Prioritise discovery and classification before expanding AI use cases. If teams cannot identify sensitive unstructured data and the AI-connected paths to it, they cannot apply meaningful access control, retention, or blast-radius limits to the systems consuming that data.
Technical breakdown
Why AI data visibility changes DSPM design
DSPM becomes a control plane for AI risk when the problem shifts from finding sensitive data at rest to understanding which systems can reach it in motion. AI applications do not just store or move information; they ingest, summarize, and recontextualize content from multiple repositories at speed. That makes visibility across SaaS, cloud, on-premises, and AI-connected stores a prerequisite for policy enforcement. In practice, the control has to show both data location and data reachability, because a dataset that is properly classified but invisible to the AI access path is still exposed.
Practical implication: map AI-connected data paths before tightening access policy, or you will govern the wrong layer.
Unstructured data classification is the weak link in AI governance
Generative AI increases exposure because it thrives on unstructured content such as text, images, and video, which traditional regex-heavy classification handles poorly. Manual tagging and pattern matching tend to produce false positives, miss context, or fail outright at file-level sensitivity decisions. That matters because AI tools often ingest exactly the content legacy systems treated as too noisy to govern well. Automated classification is not a convenience feature here. It is the mechanism that decides whether the access policy can be aligned to the actual sensitivity of what AI is allowed to consume.
Practical implication: treat file-level classification of unstructured data as a prerequisite for AI access governance.
AI tools now behave like governed non-human identities
The article's strongest identity point is that generative AI tools can be analysed as non-human identities with their own access context. That means their privileges, data reach, and blast radius need the same lifecycle attention that service accounts receive, but with a broader data security lens because the output can leak sensitive context to users and vendors. DSPM is used here to identify what those AI tools can access and to constrain that scope. The governance question becomes whether the tool's reach matches its business task, not whether it is merely authenticated.
Practical implication: inventory AI tools as non-human identities and review their effective data access, not just their login method.
Breaches seen in the wild
- 12,000 secrets in LLM training data: Truffle Security found 11,908 live API keys and passwords hard-coded in web pages captured by Common Crawl, a dataset used to train LLMs.
- Langflow Flodrix botnet 2025: Attackers used Langflow's unauthenticated code endpoint to dump AI servers' environment variables and enlist them in the Flodrix botnet.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
AI security has become a data visibility problem, not a model problem. The article is right to move the centre of gravity from the AI system itself to the data estate it can reach. In practice, the exposure path is often inherited from broad content access, not from flaws in the model. The implication is that governance teams must measure reachability before they can measure risk.
Unstructured data is where legacy data governance breaks first. Structured records were always easier to classify, but generative AI consumes the messy material that old systems struggled to tag accurately. That creates a classification debt that looks small in isolation and becomes material when AI tools can ingest it at scale. The practitioner takeaway is that file-level sensitivity has become a baseline requirement, not an edge case.
Generative AI tools now need NHI governance, not just usage policy. When an AI tool can access sensitive repositories, it behaves like a governed non-human identity with a distinct blast radius. That means access scope, data context, and offboarding all matter in the same way they do for service accounts, but the review criterion shifts to what the tool can expose through inference and output. The control gap is not authentication alone; it is unmanaged effective access.
Data reachability gap: The article identifies the core failure mode clearly: organisations can know data exists, but still not know which AI systems can touch it or what those systems can expose. That breaks the assumption that discovery and governance are separate phases. The implication is that AI-era DSPM must bind discovery, classification, and access control into one operating model.
AI governance will increasingly converge with identity governance. The most useful framing here is not a separate AI security silo, but a shared control plane for identities, entitlements, and sensitive data. As AI tools become routine consumers of enterprise content, IAM, NHI, and data security teams will have to coordinate on scope, context, and evidence. The programme implication is that access governance can no longer be evaluated without data context.
From our research library:
- Only 23% of IT leaders were very confident in their organisation's ability to manage security and governance for GenAI deployments, according to a 2025 Gartner survey of 360 IT leaders.
What this signals
Data visibility now sits at the centre of AI-era governance. Organisations do not need a separate security model for every AI use case if they can first answer a simpler question: what data can each tool actually reach? That makes DSPM a governance enabler, not just a classification utility.
The operational gap is between discovery and effective control. Many programmes can find data, but fewer can explain why a particular AI tool has access to it, or whether that access is still justified. The immediate programme risk is not theoretical exposure. It is unknown reach into content that was never meant to become machine-readable at scale.
For practitioners
- Inventory AI-connected data paths Map which SaaS, cloud, on-premises, and AI applications can reach sensitive repositories, then document the effective data scope for each path.
- Classify unstructured data at file level Replace regex-heavy, manual classification workflows with automated file-level sensitivity detection for text, images, video, and mixed-content stores.
- Treat AI tools as non-human identities Assign ownership, access reviews, and offboarding logic to generative AI tools that can read or expose enterprise content, especially where Copilot-like integrations are present.
- Constrain access by sensitivity and context Apply least-privilege data access rules based on what the AI tool actually needs, and review where vendor, external user, or cloud-provider exposure is possible.
Key takeaways
- AI security now depends on understanding what data AI tools can reach, not only where that data is stored.
- Unstructured content is the hardest part of the estate to classify, and it is exactly what generative AI consumes at scale.
- DSPM becomes more effective when teams treat AI tools as non-human identities with scoped access and explicit ownership.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 addresses the attack and risk surface, while NIST CSF 2.0 and CSA Cloud Controls Matrix set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-04 — Insecure Authentication | AI tools are treated here as governed non-human identities with access to sensitive data. |
| NHI-05 — Overprivileged NHI | The article focuses on AI tools with broader data reach than their task requires. | |
| NHI-08 — Environment Isolation | DSPM is used to limit how far AI tools can move sensitive data across environments. | |
| Recommendation — Review AI tool authentication and ownership before granting access to sensitive repositories. Reduce AI tool access to the minimum repositories needed for its business purpose. Separate high-sensitivity data domains from AI ingestion paths that do not need them. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | The core governance issue is whether AI access to data is appropriately authorised. |
| Recommendation — Apply PR.AA-05 to align AI data entitlements with least-privilege business need. | ||
| CSA Cloud Controls Matrix | IAM — Identity and Access Management | The article ties AI data exposure to identity context and access governance. |
| Recommendation — Use IAM controls to govern which AI tools can reach sensitive datasets and when. | ||
Key terms
- Data Security Posture Management: Data Security Posture Management, or DSPM, is the continuous discovery and monitoring of where sensitive data lives, how it is exposed, and where policy gaps exist. Its value rises when it feeds remediation rather than generating findings alone, especially in environments where AI expands the number of data paths.
- Unstructured Data Environment: An unstructured data environment is the collection of files, documents, shares, and content repositories that do not fit neatly into relational systems. These environments are difficult to govern because sensitive information can spread across many locations, making discovery, classification, access control, and remediation more complex.
- Non-Human Identity (NHI): A digital identity assigned to a non-human entity such as a software application, service account, API key, bot, machine, or AI agent that enables it to authenticate and interact with systems without direct human involvement. NHIs now outnumber human identities in most enterprises by 25 to 50 times.
- Blast Radius: The potential scope of damage if a specific credential or identity is compromised. Identities with broad permissions have a larger blast radius and represent a higher priority for least-privilege enforcement and security controls.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity security are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an identity or security programme, it is worth exploring.
Published by the NHIMG editorial team on June 8, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org