By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: SentraPublished August 26, 2025

TL;DR: A global travel platform secured petabytes of sensitive data across 600-plus cloud accounts in 30 days by replacing manual tagging and reactive DLP with agentless discovery, AI classification, and compliance mapping, according to Sentra. The bigger lesson is that visibility, access governance, and data classification have to move together when cloud estates outgrow manual operations.


At a glance

What this is: A Sentra case study describes how a global travel platform achieved visibility and governance across petabytes of sensitive data in more than 600 cloud accounts within 30 days.

Why it matters: It matters because identity, access, and data teams need a shared control picture when cloud estates scale beyond manual tagging, reactive alerts, and fragmented governance.

By the numbers:

👉 Read Sentra's case study on securing petabytes of cloud data across 600+ accounts


Context

Cloud data security breaks down when organisations cannot answer three basic questions quickly enough: what sensitive data exists, where it sits, and who can reach it. In this case, the cloud estate had expanded across hundreds of accounts and petabytes of storage, while manual tagging and legacy DLP kept governance slow, inconsistent, and reactive.

The primary issue is not just cloud scale. It is the control gap that appears when data classification, access visibility, and compliance mapping are managed as separate activities instead of a continuous governance process. For IAM, PAM, and data security teams, this is a familiar pattern: once the estate becomes distributed, delayed discovery turns into delayed control.


Key questions

Q: How should security teams govern cloud accounts when estates keep growing?

A: They should treat account growth as a control-design problem, not a provisioning problem. Standardise inventory, require every material change to flow through code, and make exceptions visible and reviewable. If teams cannot prove where changes came from, governance is already failing across access, compliance, and recovery.

Q: Why do manual tagging and reactive DLP fail at cloud scale?

A: Manual tagging cannot keep up with rapid storage growth, and reactive DLP only detects issues after data has moved or been exposed. At cloud scale, that means control arrives too late to reduce blast radius. The better model is continuous discovery with validated classification and clear ownership for remediation.

Q: What do IAM teams need to do differently when data visibility improves?

A: They should use new visibility to trigger access decisions, not just reporting. Once sensitive stores are identified, entitlement review, privilege reduction, and exception handling should follow quickly. That keeps data security from becoming a passive inventory exercise and turns it into a governance loop.

Q: How do compliance teams turn DSPM findings into audit value?

A: They should map sensitive data locations to the specific controls and evidence auditors ask for, then document who can access what and why. This is especially useful for PCI DSS and GDPR, where inventory, access limitation, and accountability matter. The key is to make classification output usable for control testing.


Technical breakdown

Why reactive DLP fails in large cloud estates

Reactive DLP is built to detect sensitive data movement or policy violations after content is already in flight. That model can work for narrow workflows, but it struggles when data spans hundreds of cloud accounts, multiple storage types, and fast-moving business units. The control failure is not simply weak detection. It is the absence of durable inventory, policy context, and consistent classification at the point where data is created and stored.

Practical implication: teams need continuous data discovery and classification before they can rely on alerting or DLP enforcement.

How agentless DSPM changes the control plane

Agentless DSPM reduces deployment friction by scanning cloud assets through APIs and metadata rather than installing collectors everywhere. That makes it easier to map sensitive data at scale and tie findings to compliance requirements such as PCI DSS and GDPR. In governance terms, DSPM becomes the control plane for data visibility, while IAM and access reviews determine whether the right identities can reach the right stores.

Practical implication: pair DSPM with access governance so visibility findings immediately drive entitlement review.

Why classification accuracy matters more than tagging volume

Manual tagging creates inconsistency because humans cannot keep pace with petabyte-scale change. AI-driven classification can improve coverage, but only if it is tuned to the organisation's data types, storage patterns, and regulatory needs. False positives and missed labels both weaken the programme. If the inventory is wrong, every downstream control, from audit evidence to access restriction, becomes less reliable.

Practical implication: validate classification quality as a control metric, not just the number of assets scanned.


NHI Mgmt Group analysis

Petabyte-scale data security fails when visibility is treated as a project instead of a control. Once cloud estates reach hundreds of accounts, manual tagging and periodic reviews cannot keep pace with change. The operational consequence is that teams only discover exposure after it has already widened. For identity and data governance leaders, the lesson is that inventory freshness is a control property, not an administrative task.

Cloud data estate fragmentation is the specific failure mode this case illustrates. Fragmentation appears when storage, access, and compliance evidence are split across too many accounts and too many teams for one governance model to hold together. That is where reactive DLP, inconsistent classification, and incomplete access visibility compound each other. The practitioner takeaway is to treat fragmentation as a measurable risk condition, not a reporting inconvenience.

Data visibility and identity visibility must converge. The article shows that knowing where sensitive data lives is only half the problem, because governance still depends on who can access it and under what conditions. That makes IAM, access review, and data security interdependent controls rather than separate programmes. Practitioners should align data discovery outputs with entitlement governance and audit workflows.

Compliance mapping is most useful when it shortens the distance between finding data and governing it. PCI DSS and GDPR are not solved by better charts alone. They become operationally meaningful when classification results feed access restriction, evidence collection, and exception handling. The field is moving toward continuous control alignment, and teams that keep compliance detached from runtime visibility will keep inheriting blind spots.

Agentless DSPM is becoming the practical bridge between data security and identity governance. In environments with thousands of storage locations, the value is not the tool category itself but the ability to connect sensitive data location to access decisions quickly. That is where NHI Mgmt Group sees a broader shift: governance programmes are moving from static reporting to identity-linked data control. Practitioners should design for that convergence now.

What this signals

Cloud data security programmes are moving toward continuous governance, not periodic cleanup. Once an estate spans hundreds of accounts and petabytes of storage, the practical problem is no longer whether data exists, but whether teams can keep its classification and access context current. The control expectation is shifting from retrospective evidence to live inventory, especially where PCI DSS and GDPR obligations overlap.

Data visibility is becoming an identity problem as much as a storage problem. If sensitive data discovery does not feed IAM and entitlement review, the organisation still cannot answer who can reach high-value data or whether access is justified. The emerging governance pattern is to connect discovery outputs to access decisioning, then verify that the control loop actually closes.

Cloud data estate fragmentation is the signal practitioners should watch most closely. Fragmentation creates the conditions where manual processes, delayed evidence, and inconsistent ownership all fail together, which means control design has to focus on reducing handoffs and shortening the path from finding data to governing it.


For practitioners

  • Implement continuous cloud data discovery Replace periodic scans and manual tagging with continuous discovery across cloud accounts, storage services, and regions so sensitive data is visible as the estate changes.
  • Link classification findings to access review Feed sensitive data inventories into IAM and entitlement review workflows so owners can assess who has access to high-risk stores and revoke excess access faster.
  • Measure classification quality, not just coverage Track false positives, missed labels, and time-to-detect for sensitive data so the team can judge whether classification is reliable enough for governance decisions.
  • Map data findings to PCI DSS and GDPR evidence Use data discovery outputs to generate audit-ready evidence for PCI DSS and GDPR control expectations, especially where booking data and personal identifiers coexist.

Key takeaways

  • This case shows that petabyte-scale cloud data security fails when discovery, classification, and access governance are not operationally linked.
  • The evidence points to a control gap created by hundreds of cloud accounts, more than 150K data stores, and slow manual tagging.
  • Practitioners should turn data visibility into an access and compliance workflow, not a reporting exercise.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 and GDPR define the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0PR.DS-1Sensitive data discovery and governance map directly to data security outcomes.
NIST SP 800-53 Rev 5AU-2Large-scale cloud governance depends on auditable discovery and traceable handling.
CIS Controls v8CIS-3 , Data ProtectionCIS data protection controls fit the article's focus on sensitive data visibility.
ISO/IEC 27001:2022A.5.15Access control governance is central when data visibility must inform entitlement decisions.
GDPRArt.32The article explicitly references GDPR and security of processing for personal data.

Use Art.32 to link cloud data discovery with appropriate technical and organisational measures.


Key terms

  • Cloud Data Estate: The full collection of data stores, accounts, services, and locations where an organisation keeps information in cloud environments. It includes structured and unstructured data, and governance depends on knowing where the data lives, who can access it, and how quickly that picture changes.
  • Data Security Posture Management: Data Security Posture Management, or DSPM, is the continuous discovery and monitoring of where sensitive data lives, how it is exposed, and where policy gaps exist. Its value rises when it feeds remediation rather than generating findings alone, especially in environments where AI expands the number of data paths.
  • Agentless Scanning: A scanning method that inspects workloads from outside the target environment rather than by installing software inside it. In Kubernetes, this usually means using APIs, registries, or snapshots to assess images and configurations before or around deployment.
  • Reactive DLP: Reactive DLP is a content control model that detects sensitive data issues after a policy violation or data movement has already occurred. It is useful for certain enforcement cases, but it is weaker than continuous discovery when organisations need accurate inventory and early governance.

What's in the full article

Sentra's full case study covers the implementation details this post intentionally leaves for the source:

  • How the agentless deployment was tuned across 600-plus cloud accounts and 150K-plus data stores
  • What the team changed in scanning cycles for high-memory formats and near real-time discovery
  • How the organisation reduced false positives while aligning findings to PCI DSS and GDPR
  • What the customer success and engineering collaboration looked like during rollout

👉 The full Sentra case study covers deployment details, scanning optimisation, and governance outcomes across the travel platform's cloud estate.

Deepen your knowledge

The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, identity lifecycle control, and secrets management in practical terms. It helps security practitioners connect identity controls to the broader governance demands their programmes face.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 21, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org