They should anchor the programme in continuous data discovery, classification, access visibility, and policy enforcement across cloud, SaaS, endpoints, and AI workflows. The goal is to understand where sensitive data lives, who or what can reach it, and how access changes over time. Without that baseline, security teams cannot reduce exposure, prove control effectiveness, or scale governance responsibly.
Why This Matters for Security Teams
Cloud growth and AI adoption change the security problem from protecting a few known systems to governing a moving set of data paths across cloud, SaaS, endpoints, and automated workflows. A long-term programme has to answer three questions continuously: where sensitive data is, who or what can access it, and whether those access paths still make sense after every platform change. That is why data security cannot remain a point-in-time audit exercise.
Current guidance from ISO/IEC 27002:2022 Information Security Controls and the CSA Cloud Controls Matrix both point toward ongoing control monitoring, but the real challenge is operational: discovery and enforcement must keep pace with data sprawl. NHIMG research on the Ultimate Guide to NHIs shows how fast identity and access surfaces expand once machine access enters the picture.
In practice, many security teams discover the weakest data paths only after a cloud migration, a SaaS integration, or an AI rollout has already exposed them.
How It Works in Practice
A durable programme starts with continuous data discovery, not with a single classification project. Teams should inventory sensitive data stores across cloud accounts, SaaS tenants, file systems, collaboration tools, and AI pipelines, then connect that inventory to access telemetry so they can see how data moves and who touches it. This is the baseline that makes policy enforceable rather than aspirational.
From there, classification should drive controls that are practical to maintain: encryption, tokenisation, segmentation, retention rules, masking, and conditional access. For AI-specific workflows, the security question is not only where data resides, but whether prompts, embeddings, retrieval layers, fine-tuning datasets, and output logs are holding data longer than intended. NHIMG research on the DeepSeek breach is a useful reminder that sensitive content can surface in places teams did not explicitly design as data repositories.
- Use automated discovery to identify structured and unstructured sensitive data continuously.
- Map data classes to business owners and technical controls, then review that mapping as cloud services change.
- Track both human and machine access, including service accounts, agents, and API-driven workflows.
- Enforce policy at the control points that matter most: storage, sharing, query, export, and model interaction.
- Measure drift over time so exceptions, shadow copies, and overexposed datasets are visible early.
Best practice is evolving toward policy-as-code and control evidence that can be refreshed automatically, especially when environments are changing weekly. Where AI is involved, the principle should be that data access is granted for a specific task and reviewed in context, not left open because a workflow may need it someday. These controls tend to break down when organisations centralise data without centralising ownership, because no one can reliably approve or retire access across every cloud and AI path.
Common Variations and Edge Cases
Tighter data controls often increase operational overhead, requiring organisations to balance stronger visibility against speed, developer autonomy, and AI experimentation. That tradeoff is real, especially in multi-cloud environments where teams use different storage patterns, different identity models, and different logging standards.
There is no universal standard for how far to classify every dataset, so many programmes adopt tiered classification and reserve the strictest controls for regulated, high-impact, or highly reusable data. The current guidance suggests focusing on the data that creates the greatest blast radius if exposed, including credentials, customer records, source code, and AI training or retrieval content.
Two edge cases matter most. First, data shared into AI tools may become persistent in logs, caches, embeddings, or prompt histories even when users believe it is temporary. Second, machine identities often outlive human expectations, which makes access reviews incomplete unless they include service accounts and automation. NHIMG’s 2026 Infrastructure Identity Survey shows how widespread this problem is: 67% of organisations still rely heavily on static credentials despite the risks they pose to agentic AI deployments.
For teams building long-term programmes, the goal is not perfect classification on day one. It is a repeatable operating model that can absorb cloud expansion, AI adoption, and new data flows without losing control of exposure.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 and CSA MAESTRO address the attack and risk surface, while NIST CSF 2.0, NIST AI RMF and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | ID.AM-3 | Data discovery and asset visibility are foundational to this question. |
| NIST AI RMF | GOVERN | AI adoption requires governance for data use, access, and accountability. |
| OWASP Non-Human Identity Top 10 | NHI-01 | Machine and service identities often control access to sensitive data paths. |
| CSA MAESTRO | TRM-02 | Agentic and cloud AI workflows need continuous trust and risk management. |
| NIST Zero Trust (SP 800-207) | AC-01 | Zero trust supports context-aware access across cloud, SaaS, and AI paths. |
Continuously inventory data assets and update ownership as cloud and AI environments change.
Related resources from NHI Mgmt Group
- How should security teams improve sensitive data classification across cloud and AI-driven environments?
- How should security teams govern data protection when AI adoption expands across enterprise systems and compliance obligations increase?
- How should security teams build an AI-BOM for cloud AI systems that use managed models, retrieval data, and third-party services?
- How should security teams build an AI-ready data inventory for cloud and SaaS environments?