By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: CyberhavenPublished July 29, 2026

TL;DR: The old split between data in motion and data at rest breaks down when files move from storage into GenAI tools without a traditional network event, leaving legacy DLP and posture tools blind to the full path, according to Cyberhaven. Data lineage, not state, becomes the more durable security model for AI-era data governance.


At a glance

What this is: This analysis argues that the traditional data in motion versus data at rest model no longer maps cleanly to AI-driven workflows, because the same file can move through storage, browsers, and GenAI tools without a visible boundary crossing.

Why it matters: It matters to IAM and security practitioners because AI usage expands the governance problem beyond static storage controls, forcing teams to think about lineage, access context, and sensitive-data handling across systems, users, and emerging AI workflows.

👉 Read Cyberhaven's analysis of why the data in motion vs. data at rest model breaks


Context

Data security programs have long been organised around a simple two-state model: data in motion and data at rest. That model worked when movement was observable at a network boundary and storage was easy to define, but AI tools now move data through browser interactions, prompts, and connected applications that legacy controls often do not see. For security teams, the issue is not just classification but governance across the full path of sensitive data.

The identity connection is real because access to files, repositories, SaaS apps, and AI tools determines where sensitive content can go next. Once users can paste or upload data into GenAI tools, the control problem shifts from storage-only protection to permissions, visibility, and policy enforcement across human access and AI-assisted workflows.


Key questions

Q: How should security teams prepare data access governance before enabling GenAI tools?

A: Start by reducing permission debt. Review file shares, collaboration spaces, and group-based entitlements so the AI only sees content that current business need justifies. Then verify labels, guest access, and exception paths so the model is not inheriting unmanaged exposure from the existing environment.

Q: Why do traditional data in motion and data at rest models fail for AI risk?

A: They assume data moves through a visible boundary that tools can inspect. AI workflows often use browser uploads, copy-paste, and application integrations that do not create the same boundary event, so state-based controls miss the transfer and lose the history needed for governance.

Q: What do security teams get wrong about AI and data classification?

A: They often treat classification as a labelling exercise instead of an access-control input. If sensitivity labels do not drive retrieval, sharing, and repository policy, AI can still surface protected content. Classification only matters operationally when it changes what the AI layer can see, combine, or return to a requester.

Q: How do organisations reduce AI exposure without blocking useful access?

A: Organisations should reduce exposure by removing stale data, tightening access around high-risk combinations, and restricting AI to verified datasets instead of broad repositories. That approach lowers blast radius while preserving use cases. The goal is not to stop AI access, but to make access intentional, visible, and defensible.


Technical breakdown

Why the data state model breaks in AI workflows

The two-state model assumes that data crosses a visible boundary, such as a network link into email, a file transfer, or a storage event. GenAI tools break that assumption because data can move through a browser session, a prompt, or an embedded application integration without looking like a classic transfer. That makes the distinction between motion and rest less useful for security design. The deeper issue is that the same content can exist in multiple contexts at once, copied, fragmented, and reprocessed. Security controls that depend on a single state snapshot miss that continuity.

Practical implication: design controls around data flow and provenance, not only around current location.

How legacy DLP misses AI data movement

Legacy data loss prevention tools were built to inspect well-known boundaries such as email gateways, endpoints, and network egress points. They usually rely on pattern matching, regex, or simple content inspection at the point of transfer. That approach fails when content is renamed, copied into a prompt, or transformed into something that no longer resembles the original file. It also leaves investigators without a reliable history of where the content came from or where it went. In AI workflows, the problem is not just interception. It is continuity of visibility.

Practical implication: add controls that track content across applications, not just at the perimeter.

Why data lineage is a better control model than point-in-time state

Data lineage records where sensitive data originated, how it moved, and who touched it along the way. That gives security teams a history-based view rather than a point-in-time snapshot. In AI-heavy environments, this matters because the same file may pass from a cloud repository into a chat interface, then into a downstream workflow, all within one session. Lineage creates a more durable governance model because it lets teams reason about exposure, reuse, and downstream propagation. For IAM and governance teams, lineage also helps connect file access to identity context and application usage.

Practical implication: prioritise lineage-aware telemetry where users can move sensitive content into GenAI tools.


NHI Mgmt Group analysis

Data lineage is becoming the governance layer that state-based controls were never designed to provide. The article shows why current storage-versus-transit thinking collapses once AI tools can ingest, transform, and redistribute sensitive content without a clean edge event. That creates a broader accountability problem for security and identity programmes because access decisions now influence downstream data movement. Practitioners should treat lineage as a governance requirement, not a reporting enhancement.

AI tool usage exposes a visibility gap, not just a data protection gap. Legacy DLP can still catch some obvious transfers, but it was never built to explain how content moved, who handled it, or whether it entered an AI system through copy-paste, upload, or connected integration. This is a classic control mismatch: the organisation thinks in states, while the workflow behaves in paths. Security teams need to align policy and monitoring to the path, not the snapshot.

Identity context now matters to data risk in a way many DSPM programmes underweight. If a user has permission to access a sensitive file, that permission can extend indirectly into AI-assisted workflows unless controls distinguish permitted access from permitted redistribution. That is where IAM, SaaS governance, and data security start to overlap. The practical conclusion is that file exposure should be analysed alongside user, app, and session context.

AI-era data risk is creating what can be called a data-path trust gap: organisations trust the current location of data more than the route it took to get there. That is a weak assumption when the same content can be duplicated, summarised, or repackaged by GenAI tools in seconds. The governance response has to follow movement histories, not just current classifications. Teams that keep treating storage state as the main control surface will miss the highest-risk hops.

What this signals

AI data governance is moving from point-in-time inspection to path-based control, and that shift will reshape how DSPM, DLP, and IAM programmes are measured. Teams should expect more pressure to prove not only where sensitive data sits, but who moved it, through which app, and into what kind of AI workflow.

Data-path trust gap: the next governance failure is likely to come from assuming that a file is safe because its current location is known. That assumption breaks when content is duplicated into prompts, copilots, or connected agents. Security leaders should align policy, telemetry, and identity control to the path of use, not the last known state.


For practitioners

  • Map sensitive-data paths into AI tools Trace how contracts, source code, customer records, and regulated data move from storage systems into GenAI applications, including copy-paste, uploads, and connected integrations. Focus on the routes users actually take, not just approved network egress points.
  • Add lineage-aware monitoring to DSPM Extend DSPM workflows so they record origin, movement, and downstream reuse of sensitive files across SaaS, endpoints, and AI tools. This gives analysts a timeline for investigation instead of a single overexposure snapshot.
  • Review identity permissions that enable AI data redistribution Check which users, groups, and service integrations can move sensitive content from controlled repositories into external AI tools. Separate read access from the ability to paste, upload, share, or forward that data into unmanaged environments.
  • Rework DLP boundaries for browser-era workflows Treat browser sessions, AI assistants, and collaboration apps as first-class inspection points. Legacy endpoint and gateway checks still matter, but they are not sufficient when data changes form inside the browser.

Key takeaways

  • The two-state model is too narrow for AI-era data movement because content now travels through browser and application workflows that legacy controls often cannot see.
  • Lineage is more operationally useful than point-in-time classification because it shows where data came from, how it moved, and who handled it.
  • Identity governance, DSPM, and DLP need to converge around allowed paths for sensitive content, especially when employees use GenAI tools.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, NIST SP 800-53 Rev 5, NIST AI RMF and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0PR.DS-1Data protection controls are directly challenged by AI-era movement paths.
NIST SP 800-53 Rev 5AC-6Least privilege matters when users can redistribute data into AI tools.
NIST AI RMFMANAGEAI governance requires controls around data handling and downstream risk.
NIST Zero Trust (SP 800-207)Zero trust principles support continuous verification across app and data paths.

Use the MANAGE function to establish controls for approved AI data use and ongoing monitoring.


Key terms

  • Data Lineage: The record of how data moves across systems, applications, and workflows. In security operations, lineage shows where sensitive data propagates, which identities touch it, and how a compromise could spread across connected environments.
  • Data-in-Motion: Data-in-motion is sensitive information while it is being transferred between systems, identities, or applications. For SaaS and AI programmes, the main concern is not only where data is stored, but which identities can move it, transform it, or expose it during transit.
  • Data at rest: Data at rest is information stored on disks, databases, backups, or object storage when it is not actively moving through a network or application flow. Protection usually relies on encryption, access restrictions, and strong key handling so stored information is not readable if the storage layer is exposed.
  • Lineage-Aware Monitoring: Lineage-aware monitoring tracks data movement across applications and sessions instead of inspecting only the current storage state. It is especially valuable where AI tools can ingest, transform, and redistribute sensitive content, because it preserves the path needed for governance, response, and accountability.

What's in the full article

Cyberhaven's full article covers the operational detail this post intentionally leaves for the source:

  • How the vendor's data lineage approach traces movement across endpoints, cloud apps, and AI tools
  • The specific ways browser uploads, copy-paste, and connected integrations evade legacy boundary-based DLP
  • Why the vendor argues point-in-time posture scanning cannot reconstruct the path of sensitive data
  • The example workflow for following a file back to its source system after it enters an AI assistant

👉 Cyberhaven's full post covers the lineage model, AI workflow gaps, and the control limits of legacy DLP

Deepen your knowledge

NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It is designed for practitioners who need to connect identity controls to broader security and governance programmes.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 1, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org