Data custody is the organisation’s practical ability to know where sensitive information is, who can access it, and how long it persists. It is a governance concept that becomes critical when data moves into third-party AI tools that may store or reuse inputs.
Expanded Definition
Data custody goes beyond simple data ownership or data processing. It describes whether an organisation can practically account for sensitive data across its full lifecycle, including collection, transfer, storage, access, retention, and deletion. In security and governance terms, custody is about operational control, not just legal title. That distinction matters when records move into cloud services, collaborative platforms, or third-party AI tools that may persist prompts, index content, or route data through multiple processors.
The concept is still evolving across vendors and operating models, but it aligns closely with the governance intent reflected in the NIST Cybersecurity Framework 2.0, where organisations are expected to identify assets, manage access, and maintain resilience across data handling processes. Data custody is therefore not a single control, but a test of whether data handling remains observable and enforceable after it leaves its original system.
The most common misapplication is treating data custody as a procurement issue, which occurs when teams assume contract language alone guarantees visibility, deletion, and access control after data is shared externally.
Examples and Use Cases
Implementing data custody rigorously often introduces tighter workflow controls and review overhead, requiring organisations to weigh data agility against the cost of visibility and retention enforcement.
- A legal team uploads sensitive case documents into a third-party AI assistant and later needs assurance that the content was not retained for model training or accessible to vendor staff.
- A security team labels customer records by sensitivity, then tracks where those records are replicated across SaaS platforms, backups, and analytics exports.
- An engineering group integrates an AI coding tool and must decide whether source snippets can leave controlled repositories without violating internal handling rules.
- A healthcare provider uses a document-sharing platform and needs evidence that retention settings, deletion requests, and access logs match policy and regulatory expectations.
- A finance organisation maps data custody responsibilities across business units, cloud providers, and processors to support audit readiness and reduce blind spots in NIST-aligned governance.
In each case, the key question is not only who owns the data, but who can prove where it went, who touched it, and whether it can still be removed when required.
Why It Matters for Security Teams
Security teams rely on data custody to close the gap between policy and reality. Without it, organisations can believe sensitive information is protected while it is still searchable in a vendor portal, copied into training datasets, or retained in logs long after business need has ended. That creates exposure in privacy, confidentiality, incident response, legal discovery, and AI governance.
Data custody is especially important where identity and access controls intersect with external systems. If privileged users can export data into tools that sit outside the organisation’s control boundary, traditional IAM or DLP measures may not be enough. Custody becomes a practical governance lens for deciding whether access is still enforceable after data has crossed a trust boundary. This is also why teams evaluating AI-enabled workflows should check whether a tool has storage, retention, and reuse behaviour that conflicts with internal handling rules, rather than assuming the workflow is transient.
Organisations typically encounter the custody problem only after a breach, regulatory review, or vendor dispute, at which point data custody becomes operationally unavoidable to reconstruct.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST CSF 2.0, NIST SP 800-53 Rev 5, NIST SP 800-63 and NIST AI RMF set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | GV.OV-01 | CSF 2.0 frames governance and oversight of assets and risk, which includes data custody visibility. |
| NIST SP 800-53 Rev 5 | MP-5 | Media transport and handling controls support custody across movement, storage, and disposal of data. |
| NIST SP 800-63 | Digital identity assurance underpins who may access data, which is central to custody control. | |
| OWASP Non-Human Identity Top 10 | NHI governance addresses machine identities that can move or retain data outside human oversight. | |
| NIST AI RMF | AI RMF governance addresses data management, provenance, and lifecycle risks in AI systems. |
Assign accountability for data location, access, and retention as part of governance oversight.