TL;DR: Nearly half of organisational data is sensitive or confidential, yet most organisations still lack the visibility needed to protect it as AI becomes embedded in workflows, according to Cyberhaven and IDC. Unified discovery, classification, DSPM, and DLP are becoming the practical control stack for trusted AI adoption because data context now determines both security and model quality.
NHIMG editorial — based on content published by Cyberhaven: IDC Spotlight on rethinking data security and insider risk for trusted AI adoption
By the numbers:
- 80% year over year., gh GenAI SaaS rose 80% year over year.
Questions worth separating out
Q: How should security teams govern AI adoption when data visibility is incomplete?
A: Start by treating incomplete visibility as a control weakness, not a reporting gap.
Q: Why do AI workflows make data sprawl a bigger security problem?
A: AI increases the number and speed of data retrieval paths, which means sensitive information can be copied, summarised, or exposed before traditional reviews catch up.
Q: Why do code injection flaws matter to IAM and NHI governance?
A: They matter because injected code often runs under a trusted application or pipeline identity.
Practitioner guidance
- Map sensitive data first Inventory where confidential and regulated data lives across SaaS, cloud storage, collaboration tools, and AI-connected workflows before expanding AI usage.
- Connect identity to data policy Tie IAM and NHI entitlements to data sensitivity labels so service accounts, tokens, and AI connectors are governed by what they can reach, not just by which system they authenticate to.
- Enforce DLP on AI retrieval paths Apply DLP controls to prompt inputs, file retrieval, exports, and sharing paths that AI tools can touch, then test whether policy blocks confidential data from reaching unmanaged destinations.
What's in the full article
Cyberhaven's full whitepaper covers the operational detail this post intentionally leaves for the source:
- How the IDC Spotlight frames unified discovery, classification, DSPM, and DLP as a data-centric security model
- The underlying research context behind the visibility gap affecting sensitive and confidential data
- Why data sprawl creates additional risk for AI adoption across enterprise workflows
- The specific way the paper connects trusted data foundations to trusted AI outcomes
👉 Read Cyberhaven's IDC Spotlight on data security and trusted AI adoption →
AI data sprawl and trusted AI adoption: are controls keeping up?
Explore further
Data-centric security is becoming the control plane for AI adoption. The article points to a reality many programmes still understate: AI risk is increasingly data risk, not just model risk. If organisations cannot discover and classify sensitive information quickly enough, every downstream control becomes less effective. For identity teams, that means AI access decisions must be evaluated against data sensitivity, not only user or workload entitlement. The practical conclusion is that data context now belongs inside governance design.
A question worth separating out:
Q: Which standards are most relevant for AI data security governance?
A: NIST CSF 2.0, NIST SP 800-53, and the NIST AI Risk Management Framework are the most useful starting points when AI adoption depends on sensitive data. Teams should map discovery, classification, access control, monitoring, and governance responsibilities to those frameworks, then verify that AI use cases have data-specific enforcement rather than policy statements alone.
👉 Read our full editorial: Data sprawl is outpacing trusted AI adoption controls