Subscribe to the Non-Human & AI Identity Journal

Notifications
Clear all

GenAI data leakage is exposing the gap in legacy DLP controls


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 15051
Topic starter  

TL;DR: Employees input sensitive data into AI tools once every three days on average, yet most events never trigger legacy DLP rules because copy-paste, browser uploads, embedded copilots, and personal accounts sit outside file-transfer assumptions, according to Cyberhaven Labs research. Legacy DLP is no longer enough when data moves through prompts and agentic workflows instead of discrete attachments.

NHIMG editorial — based on content published by Cyberhaven: How to Prevent Data Leakage to GenAI Applications

Questions worth separating out

Q: How should security teams stop GenAI systems from leaking sensitive data?

A: Security teams should combine runtime policy enforcement, semantic detection, and identity-aware access checks.

Q: Why do GenAI tools expose a blind spot for legacy DLP?

A: Legacy DLP was designed for file transfers, email attachments, and known upload events.

Q: What breaks when employees use personal and corporate AI accounts interchangeably?

A: Interchangeable account use breaks attribution, policy enforcement, and data handling assumptions.

Practitioner guidance

  • Instrument endpoints for prompt and clipboard visibility Deploy controls that observe copy, paste, browser uploads, and in-session AI interactions at the endpoint, because network-only monitoring will miss the majority of GenAI data movement.
  • Classify content by lineage, not only by text pattern Propagate sensitivity from the original source document into downstream prompts and uploads so a copied paragraph retains its confidential label even when the text no longer matches a pattern rule.
  • Tier AI tools by exposure and account type Build an allow, restrict, and block model that considers whether a tool trains on submitted content, how much sensitive data it receives, and how often employees use personal accounts.

What's in the full article

Cyberhaven's full blog covers the operational detail this post intentionally leaves for the source:

  • Endpoint and browser control specifics for seeing copy-paste and upload events in real time
  • Policy logic for classifying data by lineage across prompts, copilots, and agent outputs
  • Tool-tiering criteria for deciding which GenAI applications to restrict, monitor, or permit
  • Operational guidance for separating corporate and personal AI usage in enterprise environments

👉 Read Cyberhaven's analysis of preventing data leakage to GenAI applications →

GenAI data leakage is exposing the gap in legacy DLP controls?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 14635
 

Data leakage to GenAI is becoming an identity governance problem as much as a content problem. The central failure is not only that DLP was built for files, but that AI usage now spans corporate and personal identities, browser sessions, and embedded assistants inside sanctioned software. Once those paths are outside the identity control plane, visibility collapses and policy enforcement follows. Practitioners should treat GenAI access as a governed identity and data flow, not as a simple application allow list.

A question worth separating out:

Q: How should organisations govern access to data used by AI systems?

A: Treat AI data access as an identity governance problem, not just a data storage problem. Define who or what can use each dataset, what purpose is allowed, and what runtime restrictions apply. Then review humans, service accounts, and AI agents separately so entitlement scope matches actual behaviour rather than a generic AI policy.

👉 Read our full editorial: Data leakage to GenAI tools exposes a DLP blind spot



   
ReplyQuote
Share: