TL;DR: Sensitive data security now depends on discovering, classifying, and governing access across users, systems, and AI workflows, because 90% of organisations have exposed sensitive cloud data and 40% of files uploaded to generative AI tools contain PII or PCI data, according to Commvault. The practical shift is that data visibility, not perimeter control, now determines whether identity and access governance can actually reduce exposure.
NHIMG editorial — based on content published by Commvault: How Can Security Leaders Protect Their Most Sensitive Data? Data and AI security enables organizations to discover, classify, and govern access to sensitive data across users, systems, and AI solutions
By the numbers:
- 90% of organizations have exposed sensitive cloud data that can be surfaced by AI.
- 40% of files uploaded into shared with generative AI tools contains Personal Identifying Information (PII) or Payment Card Industry (PCI) data.
Questions worth separating out
Q: How should security teams govern sensitive data used by AI systems?
A: Security teams should treat AI as a data consumer that needs policy boundaries, not just authentication.
Q: Why do overpermissive accounts increase data exposure risk?
A: Overpermissive accounts widen the number of files, records, and workloads that can be reached after a compromise or mistake.
Q: How can teams tell whether data classification is actually working?
A: Look for measurable evidence that labels match reality across different data types, locations, and business contexts.
Practitioner guidance
- Expand discovery into AI data paths Inventory where sensitive content appears in training data, prompts, outputs, shared documents, and cloud repositories so discovery covers the full AI and data lifecycle.
- Bind classification to enforcement rules Map PII, PCI, PHI, secrets, and intellectual property labels to masking, redaction, retention, and sharing policies so classification changes what controls actually do.
- Review service accounts with data access Include APIs, workloads, and service accounts in access reviews, then remove unnecessary access paths that let machine identities reach regulated datasets without a current business need.
What's in the full article
Commvault's full article covers the operational detail this post intentionally leaves for the source:
- The article's stepwise view of discovery, classification, and access governance across hybrid data estates.
- Specific examples of how Commvault positions masking, redaction, and retention controls in AI workflows.
- The article's own compliance framing for GDPR, HIPAA, and PCI DSS across the data lifecycle.
- The practical description of how human and machine identities are governed together in the product context.
👉 Read Commvault's analysis of data and AI security across the full data lifecycle →
Data and AI security: why visibility and access control still fail?
Explore further
View Full Forum → | NHI Foundation Course → | Our Services →
Data classification debt is now a security liability. Organisations that know where data lives but cannot classify it accurately are operating with a false sense of control. Classification debt grows every time AI tools ingest unlabelled content, because governance rules cannot be enforced reliably against unknown data. The practical conclusion is that classification quality is now a control outcome, not a documentation exercise.
A question worth separating out:
Q: Who is accountable when AI-driven automation touches sensitive personal data?
A: The organisation remains accountable, even when access is executed by workloads, service accounts, or automated workflows. Governance must cover the identity behind the action, the data touched, and the evidence produced. If automation can access personal data, it must sit inside the same access and audit model as human users.
👉 Read our full editorial: Data and AI security depends on classification and access governance