By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: SentraPublished November 27, 2025

TL;DR: Unstructured data remains the least governed and fastest-growing enterprise data class, with the source article arguing that 80% of organisational data is effectively invisible to current security tools and volume is rising by up to 65% annually, according to Sentra. The security problem is no longer discovery alone: petabyte-scale, SaaS-spread data breaks first-generation DSPM models and forces teams to treat access, classification, and remediation as one control plane.


At a glance

What this is: This is an analysis of why unstructured data is becoming the dominant DSPM failure mode, with scale, visibility, and remediation gaps driving risk.

Why it matters: It matters because identity, access, and governance controls only work when security teams can reliably find the data, understand who can reach it, and act fast enough to contain exposure across cloud and SaaS environments.

By the numbers:

👉 Read Sentra's analysis of unstructured data security and DSPM blind spots


Context

Unstructured data is any business content that does not sit neatly in rows and columns, which makes it much harder to classify, secure, and govern at scale. In this article’s framing, the problem is not just volume but invisibility: cloud shares, collaboration tools, code repositories, and AI-generated content expand faster than legacy data security models can inspect them, creating a governance gap for security and identity teams alike.

That gap matters to IAM and data security programmes because access control is only as good as the inventory behind it. When teams cannot see where sensitive files live, who can reach them, or how sharing changes over time, identity-based policy enforcement becomes incomplete, and DSPM turns into a partial control rather than an operating model.

The article’s starting position is unfortunately typical of modern cloud-first estates, where unstructured content sprawls across SaaS and object storage faster than manual controls can keep up.


Key questions

Q: What breaks when unstructured data cannot be discovered reliably?

A: When unstructured data cannot be discovered reliably, classification and remediation become partial controls rather than enforceable controls. Security teams lose the ability to map exposure to access paths, which means oversharing, stale links, and shadow repositories persist until they are externally exposed or audited. The result is slower containment, weaker compliance evidence, and a much larger breach surface.

Q: Why do unstructured files create more governance risk than structured databases?

A: Unstructured files often lack stable metadata, consistent schemas, and predictable ownership, so the security model depends heavily on access context and manual controls. They also move easily through collaboration tools and SaaS platforms, where permissions change faster than periodic reviews can capture. That makes them harder to classify, harder to remediate, and easier to over-share.

Q: How can teams tell whether DSPM is actually improving security?

A: Teams should look for fewer unknown sensitive-data locations, faster classification of new repositories, and a tighter link between exposure findings and entitlement changes. If discovery is improving but no access decisions change, DSPM is producing visibility without governance impact.

Q: Should organisations automate remediation for sensitive unstructured data?

A: Yes, but only with guardrails. Automation should be used to revoke unsafe sharing, narrow access, and create audit trails for high-confidence findings, while ambiguous cases still route to human review. The point is to compress the time between exposure and containment without creating uncontrolled changes to legitimate business workflows.


Technical breakdown

Why unstructured data breaks traditional DSPM architectures

Traditional DSPM and DCAP tools were built around predictable repositories, stable schemas, and manageable growth rates. Unstructured data defeats those assumptions because it appears across email, chat, documents, screenshots, logs, code, and GenAI outputs, often without consistent metadata. At petabyte scale, agent-based scanning becomes too slow, too brittle, or too expensive to keep coverage current. The failure is architectural: discovery, classification, and remediation are treated as separate jobs when they now need to operate continuously across changing cloud and SaaS contexts.

Practical implication: teams need discovery and classification models that are cloud-native, continuous, and scalable rather than batch-oriented and repository-specific.

How access context changes unstructured data risk

Unstructured data is rarely dangerous only because it exists. It becomes dangerous when sharing permissions, collaboration links, and inherited access create a path from storage to exposure. That is why identity mapping matters inside DSPM: security teams need to understand which users, groups, service accounts, or external guests can reach a file set and whether that access is still justified. Without access context, classification produces labels but not containment. In practice, the governance problem sits at the intersection of data security and identity governance.

Practical implication: map unstructured-data findings back to identity and access paths before trying to remediate exposure.

Why AI-driven classification helps, but only if remediation is operationalized

The article argues that pattern matching and first-generation LLM scanning fail on natural language, images, code, and mixed-format content. AI-driven classification can improve detection, but only if it is paired with automated remediation, policy enforcement, and audit-ready reporting. Otherwise, teams simply create more alerts for already stretched analysts. The core design question is not whether AI can label data, but whether classification results can trigger safe, governed action quickly enough to reduce exposure before it spreads.

Practical implication: tie classification outputs to controlled remediation playbooks and audit trails, not to analyst queues alone.


Threat narrative

Attacker objective: The attacker objective is to locate and abuse highly sensitive unstructured content that security teams cannot reliably see or govern.

  1. Entry occurs when sensitive content lands in cloud storage, collaboration platforms, code repositories, or GenAI pipelines without adequate metadata or guardrails.
  2. Escalation follows when permissive sharing, inherited access, or stale links make that content reachable by more users than intended.
  3. Impact is realised through exposure of contracts, PII, intellectual property, or training data, often with compliance and breach costs that extend beyond the original repository.

NHI Mgmt Group analysis

Unstructured data is becoming an identity governance problem, not just a storage problem. The article is right to frame visibility as the primary failure, because access policy cannot be enforced against assets security teams cannot enumerate. In cloud and SaaS estates, file-level exposure is often driven by user, group, and guest permissions rather than by the repository itself. That makes this a governance issue for IAM, data security, and privacy teams together, not a narrow DSPM feature debate. Practitioner conclusion: treat unstructured-data discovery as an access-governance prerequisite.

DSPM 1.0 created a blind spot by separating classification from control enforcement. The article describes tools that can label or alert but cannot keep up with changing data volumes or take meaningful action. That gap matters because modern risk lives in the delta between discovery and remediation, not in discovery alone. Detection-response latency: when classification cannot trigger controlled remediation in the same operating cycle, exposures remain live long enough to become incidents. Practitioner conclusion: reduce the time between finding sensitive content and changing its access state.

The strongest control here is not better pattern matching but governed access context. The article’s most useful insight is that unstructured data becomes dangerous when its permission graph is disconnected from its sensitivity profile. That is where identity and data governance meet, and where NHI-style access patterns can matter in code repositories, automation workflows, and service-linked content stores. Frameworks such as NIST CSF and NIST SP 800-53 are relevant because they force control ownership, monitoring, and least-privilege thinking. Practitioner conclusion: align data classification with identity-aware access decisions.

Petabyte-scale unstructured data forces a shift from periodic review to continuous control. Once environments span SaaS, multi-cloud, and collaboration layers, manual governance becomes a lagging indicator rather than a control. The article’s emphasis on automated remediation is directionally correct, but only if those actions are bounded by policy and audit evidence. For practitioners, the market signal is clear: future DSPM will be judged on whether it can prove containment, not just detection. Practitioner conclusion: evaluate tools on enforceable workflow, not dashboard coverage alone.

What this signals

Detection-response latency: unstructured data security now depends on how quickly a finding can become a containment action. In cloud and SaaS environments, inventory is only useful if it is paired with access change, and that makes the operational link between DSPM and IAM more important than the classification engine alone. The NIST Cybersecurity Framework 2.0 remains a useful lens for translating this into govern, identify, protect, detect, respond, and recover activities.

The programme signal for security leaders is that data security is moving closer to identity governance. If unstructured content can be shared externally, inherited through group access, or consumed by AI systems, then data policy has to be enforced through access policy. Teams that cannot demonstrate that link will struggle to prove control effectiveness to auditors and internal risk owners.

Security leaders should also expect unstructured data to surface as an AI governance issue, not just a data-loss issue. Once training datasets, prompts, and generated content enter the same repositories as business documents, the boundary between data hygiene and model risk narrows sharply. The practical test is whether your control stack can trace exposure from file to identity to downstream system before the data is reused.


For practitioners

  • Inventory unstructured-data repositories by access path Build a cross-platform inventory for email, SaaS collaboration, object stores, code repositories, and legacy file shares, then map each repository to the identities and groups that can reach it. Use the identity map to prioritise the highest-risk sharing paths first.
  • Replace static keyword rules with context-aware classification Use classification methods that can evaluate natural language, code, images, and mixed file types, then validate outputs against high-risk repositories before relying on them for remediation decisions.
  • Automate permission fixes for exposed content Connect discovery findings to bounded remediation playbooks that can revoke public links, narrow sharing scopes, and remove stale access without waiting for manual ticket queues.
  • Tie DSPM findings to IAM and audit evidence Require every high-severity exposure to resolve into a documented access owner, a remediation action, and an audit trail that shows what changed and when.
  • Prioritise GenAI training data governance Treat unstructured content used for model training as a sensitive dataset class, then apply separate approval, lineage, and retention controls before it can influence downstream AI systems.

Key takeaways

  • Unstructured data is now the largest visibility gap in enterprise data security, and legacy DSPM designs are not built to close it.
  • The key failure is not classification alone, but the time it takes to turn discovery into enforced access change.
  • Practitioners should evaluate unstructured data controls on inventory quality, identity context, and automated remediation, not on dashboard coverage alone.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 and GDPR define the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0PR.DS-1Unstructured data exposure is a data protection and visibility problem.
NIST SP 800-53 Rev 5AC-6Over-sharing and broad file access are least-privilege failures.
CIS Controls v8CIS-5 , Account ManagementExcessive and stale access to collaboration content depends on account governance.
ISO/IEC 27001:2022A.5.15Access control policy is central to unstructured data governance.
GDPRArt.32The article’s personal-data exposure and compliance risk are directly relevant to security of processing.

Map unstructured data controls to PR.DS-1 and verify sensitive content is discovered and protected continuously.


Key terms

  • Unstructured Data Classification: The process of identifying and labelling documents, presentations, PDFs, and similar content without relying on a fixed schema. In security programmes, the goal is not just finding files, but assigning enough context for policy, access control, retention, and monitoring to work consistently across environments.
  • Data Security Posture Management: Data Security Posture Management, or DSPM, is the continuous discovery and monitoring of where sensitive data lives, how it is exposed, and where policy gaps exist. Its value rises when it feeds remediation rather than generating findings alone, especially in environments where AI expands the number of data paths.
  • Data Access Governance: Data access governance is the practice of deciding who or what should reach specific data based on sensitivity, business purpose, and observed access paths. It combines classification, entitlement analysis, and review workflows so access decisions reflect exposure, not just permission status.
  • Detection-Response Latency: The elapsed time between identifying a security issue and executing a bounded, auditable fix. In data security programmes, long latency means exposure persists after discovery, which undermines the value of detection and weakens compliance evidence.

What's in the full article

Sentra's full article covers the operational detail this post intentionally leaves for the source:

  • Agentless discovery coverage across AWS, Azure, Google Workspace, Microsoft 365, Dropbox, and legacy file shares
  • Petabyte-scale classification and risk scoring approach for unstructured repositories
  • Automated remediation playbooks for permission changes, restricted sharing, and policy enforcement
  • Deployment example showing exposed sensitive files, remediation coverage across 10 million documents, and reduced manual investigation time

👉 The full Sentra article covers the scale assumptions, remediation model, and deployment findings in more detail.

Deepen your knowledge

NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management for practitioners who need to connect access control to operational risk. It helps security and identity teams build the governance muscle that supports broader data and AI security programmes.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 20, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org