TL;DR: Data growth is accelerating toward 395ZB by 2028, while AI, SaaS sprawl and hybrid cloud expansion are multiplying the places sensitive information can reside, according to Ground Labs' analysis of IDC, Okta and Flexera data. The governance challenge is no longer just storage volume, but continuous discovery, classification and access control across fragmented estates.
NHIMG editorial — based on content published by Ground Labs: Data-centric governance in the age of AI
By the numbers:
- By 2028, worldwide data created, captured, replicated and consumed is expected to reach almost 395ZB.
- Companies with 2,000 or more employees deploy an average of 247 apps, increasing the number of data locations to govern.
- More than half of enterprise and SMB workloads now run in the cloud, and 70% operate hybrid cloud environments.
Questions worth separating out
Q: How should security teams govern sensitive data across fragmented cloud and SaaS estates?
A: Security teams should use a combined discovery and entitlement model.
Q: Why do AI and SaaS environments make PII governance harder?
A: Because the data is no longer confined to a database or a controlled application boundary.
Q: What breaks when data discovery is incomplete?
A: Risk assessment becomes guesswork because teams cannot reliably identify what sensitive data exists, where it lives or which systems can reach it.
Practitioner guidance
- Map sensitive data to identity controls Connect DSPM findings to access review, least privilege and multifactor authentication so data classification directly informs who can reach the information and through which apps.
- Prioritise shadow data and shadow AI discovery Focus first on data stores and AI usage that sit outside central security oversight, because those are the places where exposure grows fastest and visibility is weakest.
- Automate policy enforcement for high-risk data Use automated controls for encryption, tokenization, localization and deduplication where sensitive data is replicated across SaaS and cloud environments.
What's in the full article
Ground Labs' full blog post covers the operational detail this post intentionally leaves for the source:
- The article’s six-step DSPM operating model for discovery, classification, risk assessment and continuous monitoring.
- Ground Labs' interpretation of how DSPM aligns with governance and compliance frameworks across cloud and SaaS estates.
- The specific control examples for mitigation, including least privilege, multifactor authentication, tokenization and deduplication.
- The source discussion of how DSPM supports privacy legislation and information security standards in practice.
👉 Read Ground Labs' full analysis of data-centric governance in the age of AI →
DSPM and AI-driven data sprawl: what it means for governance?
Explore further
DSPM is becoming the control plane for data governance, not just a point solution for discovery. The article shows that discovery, classification, risk assessment and monitoring are now part of one operating model rather than separate workflows. That matters because data security failures increasingly emerge from fragmentation, not from a single missing control. Practitioners should treat DSPM as a governance layer that connects security, compliance and access management.
A question worth separating out:
Q: How can teams prove DSPM is working?
A: Track whether exposure is falling in priority datasets, whether classification is accurate enough to support policy decisions, and whether audit evidence can be produced without manual scrambling. Coverage alone is not sufficient. A working programme reduces risk, shortens response time, and makes compliance evidence repeatable.
👉 Read our full editorial: Data-centric governance is becoming essential as AI expands data sprawl