TL;DR: Article 10 of the EU AI Act requires evidence that high-risk AI training, validation, and test datasets are relevant, representative, and bias-screened, while most DSPM tools only answer where data sits and who can access it, according to Sentra. The gap is not legal wording but operational proof: dataset lineage, continuous evidence, and identity-aware access control now matter as much as classification.
At a glance
What this is: Article 10 of the EU AI Act turns dataset governance into an evidence problem, and Sentra argues most DSPM tools do not cover the full requirement.
Why it matters: For IAM, NHI, and security teams, this matters because AI data governance depends on knowing which identities, pipelines, and agents can reach training data, not just where sensitive data resides.
👉 Read Sentra's analysis of EU AI Act Article 10 data governance requirements
Context
EU AI Act Article 10 raises the bar for AI governance by requiring proof that training, validation, and test datasets are suitable, representative, and screened for bias. In practice, that shifts the problem from policy ownership to operational evidence, which is where many existing data security and compliance models stop short.
For security and identity teams, the important change is that dataset governance now intersects with access governance. If AI systems, data pipelines, and non-human identities can reach training data, the organisation needs controls that connect lineage, classification, and identity evidence in one audit trail.
Key questions
Q: How should security teams govern access to AI training data?
A: Security teams should treat AI training data as a privileged asset and apply least privilege, ownership, and review cycles to every identity that can read, export, or transform it. The focus should be on the pipelines that create model behaviour, not just the model runtime. If data access is broad, the AI programme inherits unnecessary exposure.
Q: Why do DSPM tools fall short of Article 10 compliance?
A: DSPM tools are usually designed to find sensitive data, classify it, and report exposure. Article 10 needs something broader: proof that the dataset is representative, cleaned, bias-screened, and traceable for the exact model it supports. If a tool cannot connect data quality to model lineage and identity access, it cannot answer the regulator’s question.
Q: What breaks when AI agent data access is not tied to identity governance?
A: What breaks is accountability. If access is not tied to identity governance, security teams cannot tell whether the agent was over-entitled, whether a policy failed, or whether the data movement was expected. That makes incident response, compliance evidence, and entitlement review much harder to perform with confidence.
Q: Who is accountable when governance fails in an AI data programme?
A: Accountability should sit with the business owner of the data domain and the control owner for the policy layer, not with a platform team alone. If stewardship, access, and quality responsibilities are not explicitly assigned, governance becomes a shared problem that no one can close.
Technical breakdown
Why dataset lineage matters more than point-in-time scanning
Article 10 is not satisfied by a snapshot of sensitive data locations. A point-in-time scan can show where data sat on a given day, but it cannot prove whether the dataset behind a specific AI system was representative, cleaned, and suitable for its intended use. Dataset lineage ties the data back to its source, transformation steps, and downstream model, which is what makes evidence auditable. Without lineage, quality claims stay generic and impossible to test.
Practical implication: map every high-risk AI dataset to its source, transformations, and model use before you treat any scan report as compliance evidence.
How identity-aware access changes AI data governance
AI data governance is not only about data quality. It also depends on which human users, service accounts, pipelines, and AI agents can access the dataset during preparation, training, and validation. That access layer matters because unauthorised or over-broad reach undermines the credibility of any governance claim. In identity terms, the relevant question is not just who owns the dataset, but which identities can touch it and whether that access is current, necessary, and reviewable.
Practical implication: extend access reviews to non-human identities and AI pipelines that interact with training data, not only to human users.
Why DSPM and Article 10 answer different questions
DSPM was built primarily to find sensitive data, classify it, and reduce exposure. Article 10 asks whether the dataset is relevant, representative, free of known defects, and bias-screened for the specific AI use case. Those are related but different control objectives. A tool can be very good at discovering PII in storage and still fail to produce the quality, context, and evidence package regulators need for a high-risk AI system.
Practical implication: pair DSPM with dataset-quality controls and evidence generation rather than assuming classification alone satisfies AI governance.
NHI Mgmt Group analysis
Article 10 creates a dataset-evidence gap, not just a compliance deadline. The practical challenge is that organisations can often describe data governance, but cannot yet prove it for a specific training set on demand. That shifts the burden from policy documentation to auditable operational records, which is a materially harder control problem. Practitioners should treat this as an evidence architecture issue, not a legal wording exercise.
Dataset lineage is becoming a governance control, not a data-management extra. Once a model depends on a specific training or validation dataset, the ability to explain where that data came from, how it changed, and who touched it becomes central to AI assurance. This is where identity governance intersects with AI governance, because non-human identities increasingly mediate access to the data that shapes model outcomes. Practitioners need lineage that can survive audit scrutiny.
Identity-aware AI data controls are now part of the AI governance stack. Article 10 indirectly exposes a familiar identity problem: access that is not continuously mapped cannot be continuously defended. If pipelines, service accounts, and agents can reach data without lifecycle oversight, the organisation loses confidence in both data integrity and accountability. Teams should align AI data readiness with identity governance, not leave them as separate programmes.
Continuous readiness is the new baseline for high-risk AI programmes. The article describes a shift from periodic review to ongoing evidence generation, which is the right direction for regulated AI. Static governance artefacts age quickly once datasets refresh, retraining begins, or agents start pulling in new data. Practitioners should assume that auditability must be designed into the operating model, not assembled after the fact.
What this signals
Dataset governance is now inseparable from identity governance when AI pipelines and agents can reach regulated data. Security teams should expect Article 10 style requirements to push more evidence work into identity, access, and lineage tooling, especially where non-human identities mediate access to training data. The practical signal is simple: if you cannot prove who and what touched the data, you cannot prove the dataset was governed.
AI governance programmes will need a control layer between discovery and audit evidence. Discovery tells you what exists, but Article 10 asks for proof of suitability, representativeness, and bias review over time. That means teams should align data discovery with lifecycle evidence and use the NIST Cybersecurity Framework 2.0 and NIST SP 800-53 Rev 5 Security and Privacy Controls to formalise accountability and review.
Dataset governance debt: this is the accumulation of missing lineage, missing access history, and missing evidence for AI training data. Organisations that delay will pay the cost later in manual audit work, unclear ownership, and controls that cannot be defended under regulatory scrutiny.
For practitioners
- Build dataset lineage into AI governance workflows Track each high-risk dataset from source through preparation, training, validation, and deployment so you can produce evidence for the exact system in scope. Treat lineage as a required control, not an optional document.
- Extend access reviews to non-human identities Include service accounts, orchestration pipelines, and AI agents in reviews for training and validation data. Remove broad or stale access that cannot be justified for the dataset’s current purpose.
- Separate discovery from compliance evidence Use discovery and classification to find data, but add dataset-quality checks, bias review records, and change history so the evidence package answers Article 10 specifically.
- Automate audit-ready evidence generation Create evidence artifacts automatically as datasets change, rather than rebuilding them manually when Legal or Audit asks for proof. The goal is a current record, not a three-week scramble.
Key takeaways
- Article 10 shifts AI governance from policy statements to dataset-specific proof that can be tested, audited, and challenged.
- DSPM remains useful for discovery, but it does not by itself prove representativeness, bias screening, or lineage for regulated AI datasets.
- Security and identity teams need continuous evidence generation around data access, pipeline identity, and dataset change history before the audit cycle begins.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST AI RMF, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the technical controls, while GDPR define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | GOVERN | Article 10 is an AI governance and accountability problem. |
| NIST CSF 2.0 | GV.OC-01 | The article centres on organisational accountability for regulated AI data. |
| NIST SP 800-53 Rev 5 | AU-2 | Audit evidence generation is central to proving dataset governance. |
| GDPR | Art.5 | The article notes overlap with GDPR-style evidence and data handling expectations. |
Log dataset changes, preparation steps, and access activity so Article 10 evidence is available on demand.
Key terms
- Data Lineage: The record of how data moves across systems, applications, and workflows. In security operations, lineage shows where sensitive data propagates, which identities touch it, and how a compromise could spread across connected environments.
- AI Data Readiness: AI data readiness is the operational state in which the data feeding an AI system is discoverable, quality-checked, access-controlled, and evidence-backed. It goes beyond storage classification by proving the dataset is suitable for the specific model, use case, and regulatory requirement.
- Identity-Aware Access: Identity-aware access is an authorization model that evaluates who or what is making a request, what it is trying to reach, and under what context. It replaces broad, persistent trust with request-level decisions. In agentic environments, it is the control that can contain a deceived agent before it reaches enterprise systems.
- Dataset Governance Debt: Dataset governance debt is the accumulated gap between what an organisation says it knows about AI data and what it can actually prove. It builds when lineage, quality evidence, and access history are not maintained continuously, making audits slow, uncertain, and expensive.
What's in the full article
Sentra's full article covers the operational detail this post intentionally leaves for the source:
- A clearer walkthrough of how its AI data readiness model maps to Article 10 evidence requirements across training, validation, and test datasets.
- Operational detail on how the platform tracks dataset lineage and access relationships inside cloud and AI data environments.
- Examples of how sensitive, stale, or redundant data is identified before it enters AI workflows.
- The article's own explanation of how these controls support audit readiness for regulated AI use cases.
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It helps practitioners connect identity controls to the operational evidence their programmes need.
Published by the NHIMG editorial team on August 18, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org