A linkage attack combines two or more datasets to identify a person or reveal sensitive facts that were not obvious in any single source. The method relies on overlapping attributes such as age, location, timestamps, or public records. It is one of the most common ways supposedly anonymised data becomes reidentifiable.
What a linkage attack does
A linkage attack works by combining data sources that each look harmless on their own, then matching overlapping attributes to reveal who a person is or expose facts that were not visible in any single dataset.
It exploits the reality that pseudonymised or “anonymised” data often still carries enough structure, such as age bands, ZIP codes, timestamps, device events, or public records, to be joined back together.
Why linkage attacks work
The attack succeeds because datasets rarely exist in isolation. One source may omit names, but another may contain a location, time pattern, role, or transaction detail that closes the gap when the two are compared.
This is especially dangerous when organisations assume that removing direct identifiers is enough. Once quasi-identifiers line up across sources, the combination can become far more revealing than either source alone.
Common places linkage attacks appear
Linkage attacks show up wherever data is shared, published, sold, or repurposed across systems. Typical examples include research releases, analytics extracts, marketing files, event logs, mobility records, and breach data combined with public information.
They are also common in environments that reuse stable identifiers or expose repeated behavioural patterns, because consistent fields make cross-dataset matching easier and more reliable.
How to think about the security impact
The main risk is reidentification, but the damage often goes further. Once separate datasets are linked, an attacker or analyst can infer sensitive attributes, reconstruct behaviour, and turn apparently low-risk data into something personal or operationally sensitive.
For that reason, the security question is not only whether one dataset is anonymous, but whether it remains anonymous after it is combined with other available sources.
Risk and Threat Considerations
Linkage attacks are a major privacy and exposure problem because they turn partial visibility into full identification. Even when one dataset looks benign, overlap with another source can expose hidden relationships, sensitive traits, or entire activity patterns.
Failure mechanism: The attacker matches quasi-identifiers, repeated timestamps, geography, demographic signals, or public records across datasets until a unique person or record emerges.
Impact: Reidentification can enable privacy harm, doxxing, profiling, discrimination, or further abuse of the linked data for fraud, extortion, or targeted exploitation.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST SP 800-53 Rev 5, NIST Privacy Framework and NIST CSF 2.0 set the technical controls, while GDPR and ISO/IEC 27001:2022 define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| GDPR | Art.25 — Data protection by design and by default | Linkage attacks directly challenge anonymous-data assumptions and privacy-by-design controls. |
| Art.32 — Security of processing | Linkage attacks expose confidentiality failures when shared data can be recombined into sensitive facts. | |
| Recommendation — Design data releases to minimise joinable attributes and prevent reidentification by combination. Apply technical and organisational measures that reduce reidentification risk in stored and shared data. | ||
| NIST SP 800-53 Rev 5 | PT-2 — Authority and Purpose | Linkage attacks become more likely when data is reused beyond the original collection purpose. |
| PT-4 — Consent | Linkage attacks matter where individuals must understand and accept how data may be combined. | |
| PT-3 — Personally Identifiable Information Processing and Transparency | Linkage attacks are a transparency problem because combined datasets can reveal more than each source alone. | |
| Recommendation — Limit secondary use of data and document the permitted purpose before sharing or combining datasets. Make combination and reidentification risks clear in notices and consent flows where applicable. Disclose how data may be matched, shared, or inferred across systems and partners. | ||
| NIST Privacy Framework | Core privacy risk management functions | The term centers on privacy risk from reidentification through data combination. |
| Recommendation — Use privacy risk management to evaluate whether released data remains safe after linkage with other sources. | ||
| NIST CSF 2.0 | PR.DS-01 — Data-at-rest is protected | Protecting data is necessary because linkage often starts with datasets that are shared, copied, or extracted. |
| Recommendation — Protect sensitive datasets in storage and limit the exposure of fields that enable cross-source matching. | ||
| ISO/IEC 27001:2022 | A.5.34 — Privacy and protection of PII | Linkage attacks are a privacy and PII-protection issue when combined datasets reveal identity or sensitive facts. |
| Recommendation — Classify joinable data as privacy-sensitive and control releases that could enable reidentification. | ||
Practitioner Guidance
Common misunderstanding: Removing names, emails, or account numbers does not by itself make a dataset safe to share. If stable attributes remain, the data may still be linkable against another source.
What to watch for: Treat datasets as linkable whenever they share age, location, timestamps, rare events, device patterns, or other quasi-identifiers. The more records can be joined across time or sources, the stronger the reidentification risk.
Practitioner takeaway: Assess anonymisation as a combination problem, not a single-file problem, and assume that external data may already exist to complete the match.