Join our Newsletter — 33% off our NHI Course
Home› Glossary› Governance, Ownership & Risk› Do Not Train Label
Governance, Ownership & Risk

Do Not Train Label

← Back to Glossary
By NHI Mgmt Group Updated September 25, 2026 Domain: Governance, Ownership & Risk

A do not train label marks data that must not be incorporated into machine learning training pipelines. It is used to prevent prohibited or consent-restricted information from being reused for model development, which helps organisations align AI practices with privacy and data usage rules.

What a do not train label is

A do not train label is a data control tag that tells downstream systems not to include specific records, documents, or fields in model training. It exists to prevent prohibited, sensitive, or consent-restricted material from becoming training input.

At its core, the label is a governance signal. It does not change the data itself, but it changes how data pipelines, dataset assembly, and model-development workflows are supposed to treat that data.

Why organisations use it

Do not train labels are used when reuse of data for model development would conflict with privacy commitments, contractual limits, internal policy, or data-use restrictions. They help organisations separate data that may be used for operations or analysis from data that must be excluded from training.

This matters because training sets are often built by aggregating content from many sources. Without a reliable exclusion marker, data can be copied into model corpora long after the original business or legal decision said it should not be reused.

In practice, the label supports data minimisation and purpose limitation by making the intended usage explicit. It is most effective when paired with clear lineage, classification, and enforcement in the pipeline that consumes the label.

How it works in machine learning pipelines

A do not train label is only useful if downstream tooling recognises and respects it. That can mean filtering at ingestion, excluding rows or files during dataset generation, or preventing a labeling review process from promoting the item into a training corpus.

The label is not the same as deletion, encryption, or access control. It is a usage constraint. A system can still store the data, display it, or process it for other approved purposes while blocking its inclusion in training.

Its value depends on consistency. If one pipeline honours the label and another ignores it, the organisation can still create training exposure from data that was supposed to remain out of scope.

This term sits at the intersection of data governance and AI governance. It is often used to operationalise a policy decision about what may be learned from, not merely what may be stored or accessed.

Do not train labels are especially important where personal data, regulated data, or customer content may enter model-development environments. They create a practical boundary between permissible operational use and prohibited training use.

Because the label depends on enforcement, it is also a model-governance issue. Organisations need confidence that excluded data stays excluded across retraining, fine-tuning, evaluation corpus creation, and vendor-managed training workflows.

Risk and Threat Considerations

When do not train controls are weak, sensitive or restricted data can be silently absorbed into training sets and then reflected in model behaviour. That creates privacy exposure, consent violations, retention problems, and the risk that restricted information becomes difficult to remove later.

Failure mechanism: The label is not read consistently, is stripped during preprocessing, or is ignored by a downstream system that builds or refreshes training data from broad source collections.

Impact: Prohibited data can influence model outputs, increase exposure of sensitive content, and create governance or regulatory findings if the organisation cannot demonstrate that excluded records were truly kept out of training.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST AI RMF and NIST SP 800-53 Rev 5 set the technical controls, while GDPR and ISO/IEC 27001:2022 define the regulatory obligations.

FrameworkControl / ReferenceRelevance
GDPRArt.5 — Principles relating to processing of personal dataDo not train labels enforce purpose limitation and minimisation for personal data.
Art.25 — Data protection by design and by defaultTraining exclusion labels are a by-design control for limiting model-use of data.
Art.32 — Security of processingReliable exclusion of restricted data depends on protected processing and controlled handling.
Recommendation — Ensure excluded personal data is not reused for training beyond the stated purpose. Build do not train enforcement into data pipelines and default training workflows. Protect pipeline handling so restricted data cannot slip into training datasets.
NIST AI RMFGovern map measure manageDo not train labels support AI risk governance over dataset scope and data usage.
Recommendation — Govern dataset usage rules and measure whether training exclusions are actually enforced.
NIST SP 800-53 Rev 5AC-3 — Access EnforcementTraining exclusion depends on enforcing who or what can use data for model development.
AU-9 — Protection of Audit InformationEvidence is needed to prove excluded data was not used for training.
Recommendation — Enforce training-use restrictions in the systems that prepare and consume datasets. Preserve logs showing excluded records were blocked from training pipelines.
ISO/IEC 27001:2022A.5.12 — Classification of informationDo not train labels are a classification treatment for controlling permitted data use.
A.8.10 — Information deletionSome do-not-train decisions align with removing data from training datasets and derivatives.
Recommendation — Classify data so downstream AI workflows can respect training restrictions. Remove prohibited data from training stores and derived corpora where required.

Practitioner Guidance

Why practitioners should care: The label only works when it is treated as an enforced control, not a courtesy marker. Teams should make sure the rule is preserved across ingestion, transformation, curation, and any external training service that receives the data.

Governance implication: Ownership should be explicit, because the most common failure is not a single technical bug but an unclear handoff between data stewards, ML engineers, and platform teams.

Practitioner takeaway: If you cannot show where the label is checked and how excluded data is blocked, the control is informational rather than protective.

Deepen Your Knowledge

Sign up to our weekly newsletter — get 33% off our NHI Foundation Level Course

    NHIMG Editorial Note
    Reviewed and updated by the NHIMG editorial team on September 25, 2026.
    NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org