Anonymous data is no longer linked to an identified or identifiable person, taking into account the means reasonably likely to be used for reidentification. Pseudonymised data still counts as personal data because the extra information needed to restore identity exists somewhere, even if it is separately stored. The practical difference is whether reidentification remains realistically possible for the recipient in the relevant context.
When does data stop being personal data?
The legal hinge is not the label on the dataset, but whether the information can still be tied back to a person using means that are reasonably likely in the real world. Once reidentification is no longer realistic for the recipient, the data is treated differently from data that still carries a practical route back to identity. That distinction drives both the compliance posture and the controls you need.
Under GDPR, the question is context-sensitive: the same record can be anonymous for one party and still be personal data for another if one party has access to additional means of identification. That is why assessments usually consider the recipient, the environment, the available extra information, and the effort needed to reverse the protection.
For teams handling regulated datasets, the boundary matters because anonymisation can remove a dataset from much of GDPR, while pseudonymisation usually reduces risk without removing the data from the GDPR regime. The difference is practical, not cosmetic, and it affects sharing, analytics, retention, and whether a dataset can be treated as outside personal-data controls.
Why pseudonymisation still stays inside GDPR
Pseudonymisation breaks the obvious link to an individual, but it does not erase the relationship between the record and the person. If a separate key, lookup table, token mapping, or other additional information exists that can restore identity, the data remains personal data because the linkage still exists somewhere in the system or process. The recipient may be protected from direct identification, but the regulator still sees a reversible personal-data relationship.
This is why pseudonymised data is best understood as risk-reduced data, not anonymous data. It can support minimisation, compartmentalisation, and safer sharing, but it does not by itself eliminate the obligations that follow from personal data. The GDPR text is explicit that security, privacy-by-design, and the context of reidentification all matter to how the data is treated.
In practice, pseudonymisation is often valuable precisely because it lets organisations separate duties: one system can work with the pseudonymised dataset, while a narrower, better controlled process keeps the reidentification mechanism. That split lowers exposure, but it only changes the risk profile if the linking material is genuinely isolated and well governed.
What anonymous data changes in practice
Anonymous data is no longer about identifying a person at all. If the dataset has been transformed so that no person is reasonably identifiable, then the main GDPR obligations attached to personal data no longer apply in the same way. This is much harder to achieve than many teams assume, because the test is not whether direct identifiers were removed, but whether identification is still reasonably possible by the intended recipient or another likely party.
That means anonymisation has to be assessed against the full context of use. Small datasets, rare attributes, linkage with public information, and cross-dataset correlation can all undermine claims of anonymity. When anonymisation is robust, the compliance benefit is significant, but the burden of proof is also higher than for pseudonymisation.
The practical takeaway is that anonymity is a stronger privacy state, but it is also a stricter standard. If there is still a realistic path back to the person, even indirectly, the safer assumption is that the data remains personal and should be governed accordingly.
Risk and Threat Considerations
The main risk is misclassifying reversible data as anonymous, then sharing or repurposing it as if it were outside GDPR. If reidentification is possible through auxiliary data, stable identifiers, or linkage attacks, the organisation can expose personal data without applying the controls, notices, retention limits, and governance expected for that class of information.
Failure mechanism: Teams remove direct identifiers but leave enough quasi-identifiers, mappings, or external join points that an intended recipient, partner, or attacker can reconstruct identity with reasonably available means.
Impact: The dataset may still be personal data, which means unlawful disclosure, weak minimisation, and incorrect regulatory treatment can create compliance exposure and real privacy harm.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5 sets the technical controls, while GDPR defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| GDPR | Art.5 — Principles relating to processing of personal data | Defines the processing principles that hinge on whether data remains personal. |
| Art.25 — Data protection by design and by default | Supports designing anonymisation and pseudonymisation into processing decisions. | |
| Art.32 — Security of processing | Relevant because pseudonymisation is a security measure that reduces exposure without removing GDPR scope. | |
| Recommendation — Assess whether the dataset still permits identification before applying personal-data processing rules. Build privacy-by-design measures that reduce identifiability at collection and sharing time. Use pseudonymisation as a security control when full anonymisation is not achievable. | ||
| NIST SP 800-53 Rev 5 | PT-2 — Authority and Purpose | Helps ensure collected data is limited to a defined purpose before anonymisation decisions. |
| DM-1 — Data Minimization | Directly supports reducing identifiability and limiting unnecessary personal data exposure. | |
| Recommendation — Document the purpose for retaining or transforming personal data before processing it further. Minimise collected attributes before attempting anonymisation or pseudonymisation. | ||
Practitioner Guidance
What to verify: Test anonymity from the recipient’s point of view, not just from the sender’s data model. If any additional dataset, key, or operational process can restore identity in a realistic way, treat the output as pseudonymised rather than anonymous.
Decision rule: If the data will be shared, analysed, or retained beyond the original context, assume pseudonymisation unless you can defend that reidentification is no longer reasonably likely. If you cannot defend that position, keep the dataset inside your personal-data controls.
Practitioner takeaway: The useful distinction is not “masked versus unmasked”, it is whether the organisation has truly eliminated the practical route back to a person for the intended recipient.
Related resources from NHI Mgmt Group
- What is the difference between a data controller and a data processor under GDPR?
- What is the difference between mapping personal data categories and documenting processing purposes under GDPR?
- What is the difference between a Data Protection Impact Assessment and a lighter assessment under UK GDPR reforms?
- What is the difference between GDPR, CCPA, and LGPD in how they treat personal data and anonymous data?