TL;DR: Healthcare privacy protections now lag far behind the ways PHI is collected, inferred, and shared across apps, wearables, data brokers, and AI systems, according to Ground Labs, leaving many organisations without durable control over sensitive health data. The governance gap is no longer limited to HIPAA coverage; it is a discovery, classification, and minimisation problem.
At a glance
What this is: This is an analysis of why US healthcare privacy law no longer covers the full PHI footprint and how that gap exposes sensitive health data across consumer and digital ecosystems.
Why it matters: It matters because IAM, data security, and governance teams need to know where health-related data and inferred attributes live so access, segmentation, and retention controls can be applied consistently.
By the numbers:
- In 2025, HHS-reported incidents averaged more than 125,000 individuals’ health records breached every single day.
- In 2022, more than a third of the top 100 US hospitals were found using the Meta Pixel tracking pixel.
👉 Read Ground Labs' analysis of the healthcare privacy crisis and PHI governance gaps
Context
Healthcare privacy now extends far beyond the traditional medical record. PHI is created, inferred, shared, and resold across apps, wearables, genetic testing services, data brokers, and adtech systems, while legacy laws were built for a much narrower model of digital healthcare data. The primary gap is not just compliance coverage, but the lack of continuous visibility into where sensitive health information actually resides and how it moves.
That creates an identity and governance problem as well as a privacy problem. Access control, segmentation, and data minimisation depend on knowing which systems, users, vendors, and analytics pipelines can touch PHI or derive health-related inferences from it. In this context, weak discovery and classification leave both human and non-human access paths ungoverned, and that is now a common rather than exceptional starting position.
Key questions
Q: How should organisations protect health data that sits outside HIPAA scope?
A: They should treat sensitivity as the organising principle, not regulatory coverage. Build a complete inventory of where health data appears, then apply classification, minimisation, segmentation, and retention controls across consumer apps, vendors, analytics platforms, and AI systems. If the data can reveal health status, it needs governance even when HIPAA does not apply.
Q: Why do inferred health attributes create privacy risk?
A: Because they can expose medical, behavioural, or demographic information without looking like clinical records. Once an organisation derives a health-related insight from non-health data, that inference can be shared, sold, or queried like ordinary profile data unless policy explicitly protects it. The control gap is usually classification, not collection.
Q: What breaks when PHI disclosure tracking is incomplete?
A: When disclosure tracking is incomplete, organisations lose the ability to explain who accessed sensitive information, for what purpose, and under what authority. That weakens incident response, patient transparency, and audit readiness. It also makes it harder to prove that access stayed within HIPAA’s permitted use and disclosure boundaries.
Q: Who is accountable when sensitive health data is exposed through vendors or AI systems?
A: Accountability should remain with the organisation that collects, processes, or benefits from the data, even if a vendor or model handles it operationally. Privacy laws may differ by state or sector, but governance responsibility does not disappear when data moves into external systems. Ownership must be explicit in contracts, controls, and review processes.
Technical breakdown
Why HIPAA leaves modern PHI exposure unaddressed
HIPAA was designed for a narrower era of healthcare data exchange, when the primary concern was protecting traditional medical records inside regulated entities and their business associates. Modern PHI now flows through consumer services, data brokers, adtech, and AI-enabled analytics that may fall outside those legal boundaries. The result is not total absence of protection, but fragmented protection that depends on the business model and data path rather than the sensitivity of the information itself.
Practical implication: teams need a data inventory that spans regulated and unregulated systems, not a compliance list limited to covered entities.
How inferred health data becomes a governance blind spot
Health-related inferences are often built from non-health signals such as location, shopping patterns, device telemetry, or model-generated classifications. Once those inferences are treated as ordinary profile data, they can be copied into downstream systems, used for targeting, or combined with other identifiers in ways that make the original sensitivity harder to recognise. This is a governance failure because the risk sits in the derived attribute, not only the source field.
Practical implication: classification rules must include inferred attributes and model outputs, not just obvious PHI fields.
Why discovery and segmentation are now control primitives
Discovery answers what data exists and where it resides, while segmentation limits who can see it and under what conditions. In privacy-heavy environments, those two controls are foundational because retention, minimisation, masking, and access review all depend on them. Without accurate discovery, organisations cannot reliably separate high-risk data from operational records, and without segmentation they cannot reduce blast radius when access is abused or systems are repurposed.
Practical implication: prioritise continuous discovery and segmented access paths before layering on policy claims or privacy notices.
NHI Mgmt Group analysis
PHI governance now depends on discovery, not just disclosure rules. The article shows that the core failure is structural: organisations cannot protect what they cannot continuously locate. Privacy programmes that rely on a static HIPAA boundary miss consumer apps, brokers, wearables, and analytics pipelines where sensitive health data accumulates. The practical conclusion is that PHI governance must start with inventory and classification across the full data estate.
Inferred health data is the new privacy edge case. Health status is increasingly reconstructed from non-health signals and AI-generated profiles, which means a record may be sensitive even when no explicit diagnosis field exists. This creates a control gap for data teams that classify only directly declared medical information. The practical conclusion is that governance must treat high-confidence inferences as sensitive by default.
Data minimisation is an identity control as much as a privacy control. When access to PHI or health-related inferences is over-broad, the issue is not only legal exposure but also unnecessary privilege across human users, vendors, and non-human services. That is especially important where vendor-managed systems and automated workflows process sensitive records at scale. The practical conclusion is to tie segmentation and least privilege to data sensitivity, not just application role.
Healthcare privacy is increasingly an AI governance problem. The article correctly notes that model training data can contain personal and health information, and that prompts can expose underlying content. That means AI systems are not neutral consumers of PHI; they can replicate, infer, and leak it unless training sets, retrieval layers, and outputs are governed. The practical conclusion is to align privacy controls with AI data handling, not treat AI as a separate risk silo.
What this signals
PHI discovery is becoming a board-level control issue: healthcare privacy failures increasingly reflect missing visibility into where sensitive data flows, not just weak notice language. Teams should expect privacy reviews to converge with data security posture management, especially where AI systems and third-party services process derived health attributes.
Health data governance now overlaps with non-human access governance: once vendors, APIs, and automated workflows can reach PHI, access scope and lifecycle control matter as much as policy wording. The practical test is whether every non-human path into sensitive health data is inventoried, reviewed, and time-bound.
Organisations should prepare for more scrutiny of inferred data, not only explicit medical records. That means classification standards, retention rules, and privacy assessments need to follow the data into consumer platforms, analytics stacks, and model pipelines.
For practitioners
- Build a cross-domain PHI inventory Map where health data, inferred health attributes, and identifiers live across EHRs, consumer apps, wearables, brokers, and AI pipelines. Include systems outside HIPAA scope because privacy exposure does not stop at regulated boundaries.
- Classify inferred health data as sensitive Extend classification rules so model outputs, enrichment fields, and profiling attributes are treated as protected when they can reveal health status or personal welfare information. This prevents derived data from being left outside governance simply because it was not collected as a clinical record.
- Segment access to high-risk PHI Separate sensitive records from general operational data and limit who and what can query them. Apply least privilege to both human users and non-human services, especially where vendor-managed systems or automation can reach large populations of records.
- Minimise retention and secondary use Delete obsolete health data, suppress unnecessary fields, and stop reusing PHI for profiling or advertising-adjacent purposes. Retention limits matter because stale data expands both breach exposure and downstream misuse.
Key takeaways
- Healthcare privacy gaps are being created by data sprawl, not only by legal loopholes.
- When health status can be inferred from consumer and AI systems, PHI governance must extend beyond the regulated record.
- Discovery, segmentation, and minimisation are now the practical controls that determine whether privacy claims hold up.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0, NIST SP 800-53 Rev 5 and NIST AI RMF set the technical controls, while GDPR and ISO/IEC 27001:2022 define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| GDPR | Art.32 | The article concerns sensitive health data protection and processing security. |
| NIST CSF 2.0 | PR.DS-1 | Data sensitivity and handling are central to the privacy gaps described. |
| NIST SP 800-53 Rev 5 | AC-6 | Least privilege is needed where vendors, apps, and AI systems can reach sensitive data. |
| NIST AI RMF | GOVERN | AI model training and prompt leakage are part of the article's privacy risk. |
| ISO/IEC 27001:2022 | A.5.12 | Information classification is essential to protecting PHI and inferred data. |
Restrict access to PHI and derived health attributes to the minimum necessary set of users and services.
Key terms
- Protected Health Information: Protected Health Information is any health-related data that can identify a person and is covered by HIPAA protections. In practice, PHI can flow through applications, integrations, service accounts, and cloud systems, which is why identity governance matters as much as data governance.
- Inferred Health Data: Inferred health data is information about a person’s health status derived from non-health signals such as location, browsing behaviour, purchases, or model output. It is often overlooked because it is created through analysis rather than direct collection, but it can be just as sensitive as explicit medical data.
- Data Segmentation: Data segmentation is the practice of separating sensitive information from broader datasets so access can be narrowed and misuse contained. In privacy programmes, it reduces unnecessary exposure by making sure high-risk fields, records, or attributes are not universally available to every user or system.
- Claim Minimisation: The practice of including only the identity attributes required for a specific access decision. In API security, claim minimisation reduces unnecessary data exposure, simplifies token review, and lowers the risk that broad identity context becomes a hidden authorisation dependency.
What's in the full article
Ground Labs' full blog post covers the operational detail this post intentionally leaves for the source:
- Examples of how health data moves through consumer apps, brokers, and wearable ecosystems beyond HIPAA-covered environments.
- Operational guidance on data discovery and classification for PHI, including where inferred health attributes should be captured.
- Practical steps for segmentation, minimisation, and cleanup of obsolete sensitive data in mixed regulated and unregulated estates.
- Discussion of how AI models can leak or reproduce health-related information from training data and prompts.
Deepen your knowledge
NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and identity lifecycle practices that support stronger access control. It is suitable for practitioners who need to connect identity governance with broader security and compliance programmes.
Published by the NHIMG editorial team on August 19, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org