Join our Newsletter — 33% off our NHI Course

How should organisations document training data to comply with AB 2013 for generative AI systems?

Organisations should maintain a data provenance record that traces where training and testing data came from, who owns it, how it was collected, and whether it was cleaned, modified, purchased, or licensed. They should also record the time period of collection and whether it is ongoing. That documentation supports public disclosure and helps prove how the model was built.

What AB 2013 is asking organisations to prove

AB 2013 is not just asking whether a model was trained on data, it is asking whether the organisation can reconstruct the dataset’s origin and handling. The practical standard is a provenance record: source, ownership, collection method, modifications, purchase or licence status, and the time window of collection. That record should be detailed enough to support disclosure and internal review.

For generative AI teams, the real test is traceability. If a dataset cannot be traced back to a source class, owner, and acquisition path, then the organisation cannot confidently explain what rights it had to use it or what transformations were applied before training.

A useful way to think about the requirement is that the documentation must follow the dataset through its lifecycle, not just describe the final training corpus. If the team merged multiple sources, filtered records, deduplicated entries, or refreshed the dataset over time, those steps belong in the record because they change what the model actually learned from.

What the provenance record should contain

The record should identify where each training and testing source came from, who owned or controlled it at the time of acquisition, and whether the data was collected directly, obtained from a third party, purchased, licensed, or otherwise authorised. It should also note whether the data was cleaned, normalised, redacted, augmented, or filtered before use.

That level of detail matters because “dataset” is often a shorthand for a chain of decisions. A compliant record distinguishes raw source material from the curated training set, and it makes later review easier when teams need to answer whether a specific sample, document, image, or log line remained intact or was materially changed.

Time period is another required dimension. Organisations should show when collection started and ended, and whether ingestion continued after an initial snapshot. For models built from continuously updated corpora, the record should show the version window used for each training run so the team can explain what was in scope at a given point in time.

When the source is licensed or purchased, the record should preserve the licence basis and any use restrictions that affect training, testing, redistribution, or disclosure. Where the source is internal, the record should still capture the business owner and the collection context so the organisation can demonstrate internal authority to use the material.

Why this matters for disclosure, auditability, and model governance

Documentation of training data is not only for legal review. It is the evidence layer that lets an organisation explain a model’s build history to auditors, regulators, customers, and internal governance teams. Without it, disclosure becomes vague and difficult to defend, especially if the organisation needs to answer questions about source quality, rights to use the data, or whether the data included restricted material.

A strong provenance record also helps separate policy failure from engineering failure. If a model behaves unexpectedly, the team can determine whether the root cause was source selection, collection bias, faulty filtering, or a licensing gap. That makes remediation more targeted and reduces the chance of reusing a contaminated or poorly understood dataset.

For broader generative AI governance, the documentation supports control over the AI supply chain and model development pipeline. NIST AI 600-1 GenAI Profile frames provenance, testing, and disclosure as core governance concerns for generative systems, while the NIST AI Risk Management Framework provides the governance structure for tracking those decisions across the model lifecycle. NIST AI 600-1 GenAI Profile and NIST AI Risk Management Framework both reinforce that provenance and traceability are governance controls, not optional paperwork.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST AI 600-1 and NIST AI RMF set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
NIST AI 600-1 Generative AI Profile GenAI provenance and disclosure are central to documenting training data sources and handling.
Recommendation — Record dataset provenance, transformations, and collection windows for each generative AI model release.
NIST AI RMF AI Risk Management Framework AI governance needs traceable data lineage to manage model risk and accountability.
Recommendation — Maintain auditable lineage for training data as part of AI governance and risk tracking.

Practitioner Guidance

What to prioritise: Build the provenance record at the dataset boundary, not after training is complete. The fastest way to fail AB 2013-style scrutiny is to rely on scattered tickets, file names, or vendor invoices instead of one authoritative record that ties source, owner, collection date, and transformation history together.

What to verify: Make sure every dataset used for pretraining, fine-tuning, evaluation, and safety testing has a clearly documented origin and authority to use it. If a source cannot be described in one sentence by owner, acquisition method, and permitted use, treat that as a documentation gap that needs resolution before the model is treated as fully documented.

Common mistake: Teams often document the final training set but omit the steps that changed it, such as deduplication, cleaning, augmentation, or ongoing refresh. That weakens the record because the legally relevant question is not only where the data came from, but what the organisation did to it before the model consumed it.

Practitioner takeaway: The standard is traceability with enough fidelity to defend the model’s data rights and build history, so the best record is the one that a reviewer can follow from source acquisition through transformation to the exact training window used.