TL;DR: Ungoverned ROT data is driving breach exposure, compliance risk, AI data leakage, and operational overhead, according to Securiti’s whitepaper, which argues that discovery alone is not enough without policy-driven deletion and lifecycle controls. The real issue is that data governance fails when organisations can find data but cannot defensibly act on it.
At a glance
What this is: This whitepaper argues that ROT data has become a governance problem, not just a discovery problem, because unmanaged retention amplifies security, compliance, AI exposure, and storage costs.
Why it matters: It matters because IAM, security, and privacy teams increasingly share responsibility for what data remains accessible to humans, NHI, and AI systems, and that access must now be governed as a lifecycle decision.
By the numbers:
- 72% of organisations have experienced or suspect they have experienced a breach of non-human identities.
- When AWS credentials are exposed publicly, attackers attempt access within an average of 17 minutes.
👉 Read Securiti's whitepaper on the cost of ungoverned data
Context
ROT data is redundant, obsolete, and trivial information that still sits in enterprise systems long after its business purpose has ended. The governance gap is not discovery. It is the inability to decide, enforce, and prove what should happen next across retention, restriction, archiving, or defensible deletion, especially when that data can be consumed by AI systems or exposed through over-retention.
For identity teams, the issue sits at the point where access, ownership, and data lifecycle intersect. Human users, non-human identities, and AI systems can all inherit exposure from stale data stores, which is why policy-based remediation matters as much as classification. This is a data governance post with direct implications for NHI, AI, and human IAM programmes.
Key questions
Q: How should security teams reduce ROT data risk without creating retention chaos?
A: Start by classifying data into retention classes with explicit owners, legal hold rules, and deletion triggers. Then automate disposition so unnecessary data is archived or removed through approved workflows instead of manual cleanup. The goal is defensible minimisation, not mass deletion. Governance works when the organisation can prove why data stayed or left.
Q: Why does shadow IT create risk for both human and non-human identities?
A: Because unmanaged SaaS often contains both employee access and machine-to-machine access inside the same application boundary. Human users may sign up directly, while API tokens, service accounts, and integrations may be created outside formal review. That combination makes the app estate a mixed identity surface that needs lifecycle control, not just software discovery.
Q: What do teams get wrong about data discovery and minimisation?
A: They often treat discovery as the finish line. Discovery only shows where data exists. Minimisation is the control that decides what happens next, including retention, restriction, archiving, or deletion. Without that step, organisations build inventories of risk instead of reducing it.
Q: Who should own accountability for AI data access risk?
A: Accountability should sit with the teams that own identity, data governance, and security operations together. If AI can access enterprise data, then ownership must cover entitlement design, monitoring, and incident response across the full workflow. The governance gap is not just technical, because without a named owner, no one can prove who approved or contained the access.
Technical breakdown
Why discovery does not solve ROT data risk
Discovery tells you where data exists, but it does not establish business purpose, retention status, or the disposition path. In practice, organisations often accumulate duplicate records, expired exports, and low-value content in places where policy enforcement is weakest. That creates a durable exposure surface for compliance, internal misuse, and AI ingestion. The technical failure is not lack of visibility. It is lack of an executable governance decision attached to the data object.
Practical implication: pair discovery with disposition rules that state whether data must be retained, restricted, archived, or deleted.
How policy-driven deletion changes the control model
Policy-driven deletion is a remediation pattern that removes unnecessary data according to approved retention and governance rules while preserving auditability. It shifts the control from manual cleanup to continuous enforcement, which is important because large data estates are too dynamic for periodic review alone. The model also reduces dependency on individual judgment, which is where retention exceptions and cleanup backlogs usually grow.
Practical implication: define deletion policies by data class, owner, and legal hold status, then enforce them automatically with logging.
Why DSPM must be tied to data minimization
Data Security Posture Management is most useful when it can do more than flag sensitive data. A mature DSPM control plane should connect discovery, classification, ownership context, and remediation into a single workflow so that high-risk data does not remain in place by default. Without that linkage, DSPM becomes another inventory layer instead of a governance mechanism that changes exposure outcomes.
Practical implication: measure DSPM by how much exposed data it removes or constrains, not just by how much it finds.
Threat narrative
Attacker objective: The objective is to turn unmanaged retention into a larger attack surface and a broader compliance failure than the organisation intended.
- Entry occurs when obsolete or redundant data remains accessible in live systems, backup stores, or AI-connected repositories long after its original purpose has ended.
- Escalation happens when that stale data is copied into more systems, exposed to broader internal access, or ingested into AI workflows without a clear retention decision.
- Impact follows when unnecessary records increase breach costs, regulatory exposure, storage overhead, and AI data leakage across teams that did not create the risk.
Breaches seen in the wild
- Coupang Signing Key Breach — Unrevoked signing key credentials expose 33.7 million records after employee offboarding failure at Coupang.
- McKinsey AI platform breach — McKinsey AI platform hack exposed 46M chats and sensitive data.
Read our 52 NHI Breaches Analysis report for a comprehensive view of breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
ROT data creates identity risk because access outlives purpose. When data has no clear owner, retention rule, or disposition path, it becomes available to people, service accounts, and AI systems that no longer need it. That is not just storage waste. It is identity exposure translated into data sprawl. The practitioner implication is that data minimization and access governance now have to be treated as one control surface.
Discovery-only programmes fail because they stop before governance action. Finding sensitive or obsolete data is useful, but the governance failure begins when organisations cannot enforce the next step. This is where manual review, exception drift, and backlog accumulation turn visibility into theatre. The practitioner implication is to measure how quickly discovered data is acted on, not how much is inventoried.
Policy-driven deletion is the named control concept this topic demands. It is the point where retention policy becomes executable rather than advisory. This matters because ROT data only stops compounding risk when deletion, archiving, or restriction happens at scale and with an audit trail. The practitioner implication is to build deletion into the control plane, not into ad hoc cleanup projects.
AI governance fails when privacy controls stay disconnected from data minimisation. If AI systems can reach stale or unnecessary data, the model training, retrieval, or assistant workflow inherits risk before any downstream safeguard can intervene. That makes ROT management a prerequisite for safe AI adoption, not a post-deployment cleanup task. The practitioner implication is to align AI enablement with retention enforcement before rollout.
The identity governance lesson is broader than data security alone. Human IAM, NHI governance, and AI access all depend on knowing whether access is still justified. A control set that cannot revoke irrelevant data access is incomplete, regardless of whether the consumer is a person, workload, or agent. The practitioner implication is to unify lifecycle rules across identity and data programmes.
From our research:
- 72% of organisations have experienced or suspect they have experienced a breach of non-human identities, according to The 2024 ESG Report: Managing Non-Human Identities.
- Another finding from LLMjacking: How Attackers Hijack AI Using Compromised NHIs shows attackers can attempt access within 17 minutes of public AWS credential exposure.
- The governance lesson extends forward to OWASP Agentic AI Top 10, where data access, tool use, and agent privilege all depend on minimised and current exposure.
What this signals
Policy-driven deletion will become a board-level data control, not a privacy back-office task. As AI systems consume more enterprise content, the question is no longer whether data can be found but whether it should still exist. Teams that leave ROT in place will keep funding risk they never intended to keep, which makes retention enforcement a programme metric rather than a cleanup activity.
Ungoverned data is becoming identity debt. When stale records remain accessible to users, service accounts, and AI workflows, organisations inherit ongoing exposure without a clear owner for remediation. That is why lifecycle thinking must now span data, access, and AI consumption paths rather than treating them as separate controls.
With 72% of organisations already experiencing or suspecting a breach of non-human identities according to The 2024 ESG Report: Managing Non-Human Identities, the risk model is already structural, and over-retention simply enlarges the blast radius.
For practitioners
- Define executable retention rules Map each major data class to a retention period, legal hold condition, and deletion trigger so disposition is no longer a manual judgement call.
- Link DSPM findings to remediation Require every sensitive-data discovery to land in a workflow that can restrict, archive, or delete the record with a logged approval trail.
- Stop AI systems from consuming ROT data Block low-value or expired content from vector stores, copilots, and downstream retrieval layers until the data has been minimised and reclassified.
- Align access reviews with data purpose Review whether human users, service accounts, and AI systems still need access to the same datasets after business purpose has expired.
- Track remediation, not just discovery Measure how much obsolete data was removed, constrained, or archived in each cycle, and treat unresolved ROT as a governance backlog.
Key takeaways
- ROT data is a governance failure because visibility without enforcement leaves stale information available to people, workloads, and AI systems.
- The risk is measurable across breach exposure, compliance burden, storage cost, and operational drag, which is why minimisation must be continuous rather than periodic.
- Security, privacy, and identity teams need shared ownership of retention, deletion, and access justification if they want to reduce exposure instead of cataloguing it.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 address the attack surface, NIST CSF 2.0, NIST SP 800-53 Rev 5 and NIST AI RMF set the technical controls, and ISO/IEC 27001:2022 define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-03 | The article centres on unmanaged non-human exposure and lifecycle control gaps. |
| NIST CSF 2.0 | PR.AC-4 | The post links data minimisation to least-privilege access and governance. |
| NIST SP 800-53 Rev 5 | AC-3 | Access enforcement is central when stale data remains reachable across systems. |
| ISO/IEC 27001:2022 | A.5.12 | Information classification and handling are needed to govern ROT disposition. |
| NIST AI RMF | MANAGE | AI exposure from stale data requires operational controls around lifecycle risk. |
Map ROT-related machine access to NHI-03 and remove unnecessary entitlements tied to stale data.
Key terms
- Rot Data: Redundant, obsolete, and trivial data that remains in systems after it has lost clear business value. In security terms, it becomes avoidable exposure because it still consumes storage, can be accessed, and may be ingested by AI or copied into downstream systems.
- Policy-Driven Deletion: A control pattern that removes data based on approved retention and disposition rules rather than manual cleanup. It turns governance into an executable workflow, which matters when organisations need defensible minimisation and an audit trail for why information was removed or retained.
- Data Security Posture Management: Data Security Posture Management, or DSPM, is the continuous discovery and monitoring of where sensitive data lives, how it is exposed, and where policy gaps exist. Its value rises when it feeds remediation rather than generating findings alone, especially in environments where AI expands the number of data paths.
- Claim Minimisation: The practice of including only the identity attributes required for a specific access decision. In API security, claim minimisation reduces unnecessary data exposure, simplifies token review, and lowers the risk that broad identity context becomes a hidden authorisation dependency.
What's in the full article
Securiti's full whitepaper covers the operational detail this post intentionally leaves for the source:
- A cost breakdown for breach exposure, regulatory fines, AI data leakage, and storage overhead tied to ROT data.
- A policy-driven deletion approach for organisations that need defensible minimisation with an audit trail.
- Operational guidance for using Securiti's DSPM capability to classify, contextualise, and remediate ungoverned data.
- Examples of how teams can move from discovery to continuous audit-ready posture.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity security are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an identity security programme, it is worth exploring.
Published by the NHIMG editorial team on August 18, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org