By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: Ground LabsPublished January 16, 2026

TL;DR: Sensitive data protection still fails first at discovery, because fragmented on-premises, cloud and SaaS estates hide PII in databases, collaboration tools, logs and backups, according to Ground Labs. For security and privacy teams, the hard problem is building a usable inventory before encryption, access control and compliance can work at scale.


At a glance

What this is: This blog argues that data discovery is the hardest part of data protection because organizations cannot protect sensitive data they cannot reliably locate across fragmented environments.

Why it matters: For IAM, NHI, and broader security teams, incomplete data visibility weakens access control, privacy governance, and lifecycle decisions because controls depend on knowing where sensitive data and the identities that can reach it actually live.

By the numbers:

👉 Read Ground Labs' analysis of why data discovery is the hardest part of data protection


Context

Data discovery is the control plane for everything that follows in data protection. If an organisation cannot locate sensitive data across databases, collaboration tools, file shares, logs, backups and SaaS services, then encryption, access control and compliance checks are operating with incomplete context rather than complete governance.

This problem now intersects with identity governance because data access is mediated by humans, service accounts, API keys and application integrations. Where data estates are fragmented, access rights become harder to review, sensitive records become easier to overexpose, and lifecycle controls lose accuracy across both human and non-human identities.


Key questions

Q: How should security teams implement data discovery in complex environments?

A: Start by mapping endpoints, databases, file shares, cloud services and SaaS platforms so the discovery scope matches the real estate where sensitive data actually lives. Then validate accuracy with sampling, because false positives and missed repositories undermine classification, DSPM and audit readiness. Discovery should produce an inventory that other controls can trust.

Q: Why does unstructured data create identity governance risk?

A: Unstructured data creates risk when access is spread across repositories and shares without clear entitlement ownership or review. In that model, sensitive content can remain accessible long after the business need has passed, and the IAM programme loses visibility over who can reach it and why.

Q: What breaks when discovery is missing from data protection programmes?

A: Encryption and access controls still help, but they cannot be applied consistently if teams do not know where sensitive data resides. Missing discovery leads to incomplete inventories, weak audit evidence and slow remediation. It also creates identity problems, because human and non-human access rights cannot be accurately reviewed against data that has not been mapped.

Q: How do organisations know whether discovery is working?

A: Look for measurable coverage across the full estate, including shadow IT, backups, logs and collaboration tools, plus evidence that classified findings feed into access review and remediation workflows. If discovery only finds known systems, it is not changing governance. Effective discovery reduces unknown repositories and shortens the time from identification to control action.


Technical breakdown

Why fragmented estates break data discovery

Modern enterprises distribute data across on-premises systems, multiple clouds and SaaS platforms, which removes any single authoritative inventory. Sensitive records also appear outside formal repositories, including email, collaboration tools, logs and storage backups. That means discovery is not a one-time scan but a continuous classification problem across changing storage patterns, permissions and formats. Shadow IT and dark data make the task harder because they create repositories that security teams do not fully track or govern. Practical implication: build discovery around continuous inventory rather than periodic file searches.

Practical implication: treat discovery as an always-on inventory problem, not a one-off scan.

How unstructured data and AI increase hidden data risk

Unstructured data has become the dominant data class in many enterprises, and it often contains the most sensitive content because it is easiest to create and hardest to govern. AI applications add another layer of risk by generating, transforming and redistributing data during normal business use, which expands the volume of material that must be classified and controlled. The result is a larger dark-data surface and a higher chance that sensitive information sits outside intended governance boundaries. Practical implication: align discovery controls with data creation paths, not just storage locations.

Practical implication: extend discovery into AI-enabled workflows where data is created or transformed.

What a centralized data inventory actually changes

A centralized data inventory turns discovery from an ad hoc search function into a governance prerequisite. It provides the evidence base for risk assessment, access reviews, privacy obligations and remediation prioritisation. Without that inventory, organisations can only guess where sensitive data resides or which systems create the biggest exposure. In practice, this is also where identity intersects with data governance, because permissions and usage can only be assessed once the data itself is mapped. Practical implication: use the inventory as the reference point for access, retention and remediation decisions.

Practical implication: anchor access reviews and remediation plans to a governed data inventory.


Threat narrative

Attacker objective: The attacker or accidental insider wants to locate and extract sensitive data from systems the organisation does not fully know exist or cannot fully govern.

  1. Entry occurs when sensitive data is copied into collaboration tools, backups, logs or SaaS services outside the primary system of record.
  2. Escalation follows when fragmented ownership and weak classification let users, integrations or support accounts retain broader data access than intended.
  3. Impact is the exposure of personal information, regulatory non-compliance and a wider blast radius for any account compromise or misconfiguration.

NHI Mgmt Group analysis

Data discovery is now a governance dependency, not a visibility nice-to-have. Organisations routinely frame discovery as a tooling problem, but the real issue is that every downstream control assumes the inventory is already complete. When sensitive data is scattered across clouds, SaaS and endpoint-adjacent stores, policy enforcement starts from partial knowledge. That weakens privacy governance, auditability and lifecycle decisions at the same time. The practical conclusion is simple: if the inventory is incomplete, the control environment is incomplete.

Hidden data creates hidden identity risk. Once PII and other sensitive records sit in collaboration systems, logs or backups, access review becomes less reliable because the organisation no longer knows which human and non-human identities can reach the data. This is where IAM, PAM and data governance overlap in a meaningful way. Access to data cannot be governed cleanly when the data estate itself is unresolved. Practitioners should treat discovery gaps as identity exposure multipliers.

Unstructured data is the new dark-data fault line. With unstructured content making up the majority of enterprise data, the challenge is less about storing information and more about identifying what deserves protection. AI-driven data creation compounds this by multiplying copies, derivatives and transient states. That creates an environment where classification debt accumulates faster than remediation can catch up. Teams need to assume that the risk surface will keep expanding unless discovery is embedded into business workflows.

Centralized inventory is the only scalable starting point for privacy and access governance. Regulatory pressure increasingly converges on evidence of where sensitive data resides, who can access it and how it is controlled. That means inventory quality becomes a measurable governance capability rather than an administrative output. For security leaders, the question is no longer whether discovery is helpful. It is whether the organisation can prove control without it.

What this signals

Discovery debt will increasingly become identity debt. As organisations map more of their data estate, they will uncover stale permissions, unmanaged integrations and service accounts with reach into sensitive repositories. That means data discovery programmes should be treated as input to identity governance, not a separate compliance exercise. The next maturity step is connecting the inventory to access review, rotation and offboarding controls.

The practical pressure point is automation. Manual cataloguing does not scale when unstructured data and AI-generated derivatives keep expanding, so teams will need discovery engines that feed directly into governance workflows and exception handling. For identity teams, that also means reviewing how machine identities interact with data stores, because hidden access paths are often the first place governance failures show up.


For practitioners

  • Build a centralized sensitive-data inventory Map databases, collaboration platforms, file shares, logs, backups and SaaS repositories into one governed inventory before expanding control policies.
  • Prioritise unstructured data classification Focus classification effort on documents, emails, exports and backups because those stores usually contain the highest volume of hidden PII and regulated content.
  • Link discovery to identity review workflows Tie data location results to access recertification so human users, service accounts and application integrations are reviewed against actual data exposure.
  • Extend discovery into AI workflows Track where AI tools create, transform or duplicate sensitive records so new data copies do not bypass retention and protection rules.
  • Use discovery findings to drive remediation order Rank systems by sensitivity, exposure and ownership clarity so teams fix the most consequential repositories first rather than treating all data stores equally.

Key takeaways

  • Data protection fails early when organisations cannot see where sensitive information lives across fragmented systems.
  • Discovery gaps are not only a privacy issue, they also distort access governance for human and non-human identities.
  • The control strategy that matters most is a centralized inventory linked to classification, identity review and remediation workflows.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 and GDPR define the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0ID.AM-01Discovery and inventory are central to asset understanding and data visibility.
NIST SP 800-53 Rev 5AU-6Audit review and analysis depend on knowing where sensitive data is stored and accessed.
CIS Controls v8CIS-1 , Inventory and Control of Enterprise AssetsDiscovery starts with knowing what data stores and repositories exist across the estate.
ISO/IEC 27001:2022A.5.9Information inventory supports governance over sensitive data and its lifecycle.
GDPRArt.30PII discovery supports records of processing and accountability for personal data.

Inventory sensitive-data locations continuously and feed findings into governance and access decisions.


Key terms

  • Data Discovery: Data discovery is the process of finding where information lives across cloud, SaaS, endpoints, backups, and analytics systems. In practice, it creates the inventory that makes classification, access decisions, recovery planning, and AI governance possible rather than speculative.
  • Dark Data: Data that an organisation has stored but no longer actively inventories, classifies, or governs. It may contain regulated information, credentials, or operational records, but the organisation lacks reliable visibility into where it lives, who can access it, and how long it should remain retained.
  • Unstructured Data Classification: The process of identifying and labelling documents, presentations, PDFs, and similar content without relying on a fixed schema. In security programmes, the goal is not just finding files, but assigning enough context for policy, access control, retention, and monitoring to work consistently across environments.
  • Sensitive Data Inventory: A sensitive data inventory is the live record of where protected information exists, how it moves, and which identities can access it. It is the operational foundation for enforcement, auditability, and breach response because classification without inventory cannot be verified.

What's in the full article

Ground Labs' full blog post covers the operational detail this post intentionally leaves for the source:

  • How its discovery approach maps sensitive data across databases, collaboration tools, logs, backups and SaaS repositories
  • Which 300-plus PII data types and country-specific patterns the tooling recognises during classification
  • How remediation workflows are structured once hidden personal data is found
  • What implementation teams need to consider when moving from discovery findings to governance action

👉 The full Ground Labs post expands on discovery methods, hidden data locations and remediation context.

Deepen your knowledge

The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, workload identity, secrets management and lifecycle controls. It helps security and identity practitioners connect discovery gaps to the access and lifecycle decisions that reduce exposure.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 18, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org