Join our Newsletter — 33% off our NHI Course

What is the difference between GDPR, CCPA, and LGPD in how they treat personal data and anonymous data?

GDPR and LGPD take a broader view of personal data, covering directly and indirectly identifiable information. CCPA is more limited and also includes household and device data. For anonymous data, CCPA is the most permissive, while GDPR allows anonymous data but not pseudonymous data. LGPD is broad enough that many data types remain regulated unless a narrow exception applies.

How GDPR, CCPA, and LGPD Draw the Line Around Personal and Anonymous Data

All three laws try to separate data that can identify a person from data that truly cannot, but they do it with different thresholds and legal effects. The key practical difference is how broadly each regime treats identifiability, because that determines when data remains regulated, when it can be shared more freely, and when anonymisation is strong enough to exit scope.

Under GDPR and LGPD, the question is usually whether a person can be identified directly or indirectly, including by combining data points. Under CCPA, the scope is narrower in some respects but still reaches data linked to a consumer or household. That means the same record set can fall inside one regime, remain partially regulated in another, or fall out of scope only after genuine anonymisation.

Why Anonymous Data Is Treated Differently Under Each Law

Anonymous data is the clearest dividing line, but the legal test is not just whether a name has been removed. GDPR allows genuinely anonymous data outside its scope, yet pseudonymous data still counts as personal data because reidentification remains possible. LGPD is similarly broad in practice, so organisations often need a high-confidence anonymisation standard before treating data as unregulated.

CCPA is generally more permissive toward anonymous data, but that permissiveness depends on whether the data can be linked back to a consumer or household. If linkage is still possible through device identifiers, profiles, or other correlating fields, the safer assumption is that the data remains covered. In practice, the difference is less about labels and more about reversibility, linkage risk, and whether the organisation can still reasonably reidentify the subject.

What This Means for Data Classification and Governance

For practitioners, the real task is to classify data by identifiability, not by the source system that produced it. A customer record, device identifier, behavioural profile, or hashed field may be anonymous in one dataset and identifiable when combined with other data. That is why classification rules, retention rules, sharing rules, and vendor contracts need to align with the strictest interpretation that applies to the use case.

This is also where governance breaks down most often: teams assume that masking, hashing, or removing a direct identifier automatically makes data anonymous. It does not. If the organisation still has a practical way to link the data back to a person or household, the safer and usually legally correct position is that it remains regulated personal data. For a practical governance baseline, the general GDPR text remains the clearest reference for the indirect-identifiability test, while the NIST Privacy Framework is useful for structuring data classification and privacy risk decisions.

Risk and Threat Considerations

Misclassifying anonymous, pseudonymous, and personal data creates regulatory exposure and also increases privacy and security risk. The practical failure mode is that data teams treat linkable data as if it were de-identified, then move it into analytics, sharing, or vendor environments without appropriate controls. Once linkage exists, reidentification, secondary use, and over-collection become the main sources of harm.

Failure mechanism: Weak anonymisation, poor dataset segregation, or hidden linkage fields let organisations reidentify people or households after they believe the data has left scope. The risk is highest when device data, customer attributes, or behavioural data can be joined across systems.

Impact: The organisation may breach privacy law, overexpose regulated data, and lose control over downstream sharing, retention, and subject-rights handling. At scale, the same classification error can affect entire analytics pipelines and third-party disclosures.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0 and NIST AI RMF set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
NIST CSF 2.0 PR.DS — Data Security Data classification and protection depend on whether data remains personal, pseudonymous, or anonymous.
GV.RM — Risk Management Strategy The scope decision hinges on privacy risk from linkable or reidentifiable data.
GV.OV — Oversight Governance is needed to enforce consistent privacy classification across teams and vendors.
Recommendation — Classify data by identifiability and apply protections based on reidentification risk. Set policy to treat linkable datasets as regulated until reidentification risk is acceptably low. Require oversight for data classification decisions that affect legal scope and sharing.
NIST AI RMF MAP 2.1 — Map Context Privacy context mapping helps distinguish personal, pseudonymous, and anonymous data use cases.
Recommendation — Map the data context before deciding whether a dataset is truly anonymous.

Practitioner Guidance

What to verify: Before calling data anonymous, test whether any realistic combination of fields, vendor data, or internal reference tables can relink it to a person or household. If yes, treat it as regulated data, not as a de-identified exception.

Decision rule: If the dataset can still support targeted action, persistent profiling, or subject-level correlation, assume the applicable law will treat it as personal data even if direct identifiers are absent. Reserve the “anonymous” label for data that is no longer reasonably reidentifiable in context.

Practitioner takeaway: The safest operational standard is to classify by reidentification risk, because the legal boundary is driven by whether identity can still be inferred, linked, or recovered, not just whether a direct identifier was removed.