By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: CyberhavenPublished May 25, 2026

TL;DR: Employees are submitting sensitive data into AI tools at scale, with Cyberhaven Labs finding that 39.7% of the data shared with AI tools is sensitive, exposing gaps in legacy DLP that was built for email, USB, and file transfer channels. The real issue is governance of data flow, context, and AI tool inventory, not just content inspection.


At a glance

What this is: DLP for GenAI extends data loss prevention to AI prompts, uploads, and agentic workflows, and the key finding is that employees are already sending sensitive data into AI tools at a rate traditional DLP was not built to control.

Why it matters: This matters to IAM and security teams because AI tools now sit inside identity-bound workflows, where access, data classification, and tool approval all shape whether sensitive information leaks beyond policy boundaries.

By the numbers:

👉 Read Cyberhaven's analysis of DLP for GenAI and sensitive data exposure


Context

DLP for GenAI is a data governance problem as much as a data security one. Traditional DLP assumes data moves through fixed channels such as email, file transfer, or removable media, but AI prompts, browser assistants, and agentic workflows break that model because the sensitive data is often copied, pasted, synthesised, and reused without a conventional transfer event. That makes the primary control question in AI tool governance much broader than block or allow.

The identity angle is real because AI usage is increasingly mediated by users, service accounts, browser sessions, and AI agents that act inside enterprise workflows. When those identities can submit sensitive data to external tools, the security team needs to understand not just what the tool is, but who or what is allowed to use it, what data it may touch, and how those actions are enforced across the lifecycle.


Key questions

Q: How should organisations control sensitive data in GenAI tools?

A: Organisations should treat prompts, uploads, and model outputs as governed data flows, then apply classification, inspection, and logging at the point of use. The control objective is to stop sensitive information from entering AI workflows without visibility. That requires policy, access rules, and monitoring to work together, not as separate programmes.

Q: Why do AI prompts create more data loss risk than traditional file transfers?

A: AI prompts often copy data out of its original container without creating a classic file-transfer event, so perimeter DLP misses the move. Prompts also mix context, making text-based detection noisy. That is why lineage and endpoint observability matter more than content scanning alone.

Q: What do organisations get wrong about DLP for AI use cases?

A: They assume keyword matching can distinguish legitimate work from sensitive exfiltration. In practice, AI prompts are contextual, so the same text may be safe in one workflow and dangerous in another. Teams need policy that evaluates intent, destination, and action, not just strings.

Q: When should institutions treat AI agents as identities rather than tools?

A: Institutions should treat AI agents as identities when the agent can authenticate, call APIs, move data, or take action without a person supervising each step. At that point, the agent affects access decisions and must be governed with the same ownership, logging, and revocation discipline as other non-human identities.


Technical breakdown

Why legacy DLP fails on AI prompts and browser-based submission

Legacy DLP is built around discrete transfer points, where content can be scanned as it leaves a mailbox, endpoint, or file share. GenAI usage often bypasses those checkpoints because the user copies text from a document, pastes it into a prompt, and submits it through a browser or embedded assistant. The data is no longer a file, and pattern matching alone cannot reliably reconstruct origin or intent. That is why AI data leakage needs endpoint observability plus context from the source document, not just outbound content inspection.

Practical implication: move from perimeter-only inspection to endpoint-level visibility that can see copy, paste, and browser submission events.

Data lineage and classification for AI tool governance

Data lineage means tracking where data came from, how it was transformed, and where it moves next. In GenAI use cases, lineage lets a control inherit sensitivity from the source document even when the pasted text no longer looks like a protected record. That is critical because AI prompts often contain fragments that are harmless in isolation but sensitive in context. This also maps to identity governance, because the user or agent submitting the data should inherit the policy constraints attached to the source information.

Practical implication: classify source data once, then propagate policy decisions to prompts and AI-assisted workflows through lineage-aware controls.

AI agents as data access and exfiltration paths

Agentic AI changes the control problem because the system can query documents, call APIs, and assemble outputs without a human manually copying each item. That means the data exposure path is no longer just human prompt abuse, but also automated collection and redistribution by delegated software identities. In practice, AI agents behave like privileged non-human identities when they operate inside enterprise systems, so governance must cover permissions, output destinations, and tool inventory together.

Practical implication: treat AI agents as governed identities and bind their access to scoped, auditable, policy-enforced permissions.


Threat narrative

Attacker objective: The objective is to move sensitive enterprise data into an AI environment where it can be retained, learned from, or redistributed outside the organisation's intended control boundary.

  1. Entry occurs when an employee, contractor, or AI agent accesses sensitive information and moves it into a GenAI interface through copy, paste, upload, or tool-calling.
  2. Credential or context abuse follows when the AI tool or workflow receives data outside the original security boundary and the organisation loses visibility into how that content is stored, retrained, or reused.
  3. Impact occurs when confidential code, customer records, contracts, or credentials are exposed beyond policy control, persisted in provider systems, or redistributed into downstream outputs.

NHI Mgmt Group analysis

AI data leakage is now an identity governance problem, not just a content filtering problem. Once employees and AI agents can move sensitive information through browser prompts, file uploads, and delegated workflows, the question becomes who is allowed to submit what data to which tool under which policy. Legacy DLP answers only part of that question because it was built for channels, not for identity-mediated data flows. Practitioners need governance that follows the user, the agent, and the data together.

Context-aware enforcement is the right concept for GenAI, but lineage is the real control gap. Pattern matching alone cannot distinguish a harmless prompt from a leaked customer record when the same text may be valid in one context and sensitive in another. Source-aware lineage closes that gap by tying policy to the origin of the data rather than its final form. That shifts DLP from reactive detection to policy inheritance, which is the only scalable model for prompt-based workflows.

AI agents should be treated as non-human identities with bounded data authority. When agents can query systems, summarise records, and pass results onward, they become a new class of governed data mover with privileges that outlast a single prompt. That demands identity-first controls around tool access, delegated scope, and output restrictions. The practical conclusion is that AI governance and identity governance are converging at the data boundary.

Shadow AI creates a verification trust gap that security teams cannot solve with awareness training alone. Employees are already using unsanctioned AI tools across business workflows, and the security problem is no longer whether use exists but whether it is inventoryable, policy-bound, and auditable. This is where identity and access governance meets data security, because every unmanaged tool is also an unmanaged trust decision. Teams should measure the gap between approved AI usage and observed AI usage as a governance metric.

Named concept: prompt lineage control. This is the ability to enforce data policy on AI prompts based on the origin and classification of the source information, not just on the text submitted. It matters because it turns copy-paste and agentic reuse into governable events instead of blind spots. Organisations that do not build prompt lineage control will keep treating AI leakage as a point problem when it is really a workflow problem.

What this signals

Prompt lineage control: organisations now need a control model that ties AI submission decisions back to the source classification, the active identity, and the destination tool. That is a practical extension of data security into identity governance, and it should be measured as part of AI tool approval rather than left as a local endpoint setting.

The wider programme signal is that AI adoption is becoming a governance inventory problem. Teams need to know which AI tools are used, which identities can reach them, and which data classes are permitted to flow through them, because unmanaged prompts are simply another form of unmanaged access.

Identity teams should expect DLP, IAM, and AI governance to converge around browser sessions, delegated agents, and source-aware policy enforcement. The organisations that move first will be the ones that can prove control over sensitive data movement rather than merely reporting on blocked events.


For practitioners

  • Implement lineage-aware DLP Track the origin of sensitive content so copied or pasted text inherits the classification of the source document before it reaches a GenAI tool.
  • Inventory approved AI tools and agent workflows Maintain a continuously updated inventory of AI assistants, browser features, and agentic workflows, including how each tool stores, trains on, or forwards submitted data.
  • Bind AI access to scoped identities Restrict which users, service accounts, and AI agents may submit sensitive data, then enforce policy by role, context, and destination.
  • Separate allow, warn, and block policies Use differentiated controls for general queries, sensitive data, and high-risk channels so enforcement reduces shadow usage instead of driving it underground.

Key takeaways

  • DLP for GenAI is really about controlling identity-mediated data movement, not just scanning content at the edge.
  • Legacy DLP misses the most common AI leakage paths because prompts, pastes, and agentic workflows do not behave like traditional file transfers.
  • Lineage, tool inventory, and scoped identity controls are the practical combination that can keep sensitive data out of unmanaged AI pathways.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 address the attack surface, NIST AI RMF, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the technical controls, and GDPR define the regulatory obligations.

FrameworkControl / ReferenceRelevance
NIST AI RMFGOVERNAI data governance and accountability are central to GenAI DLP decisions.
NIST CSF 2.0PR.AC-4GenAI access decisions depend on least-privilege enforcement and data handling rules.
NIST SP 800-53 Rev 5AC-6Least privilege directly governs who can submit sensitive data to AI tools.
OWASP Agentic AI Top 10Agentic workflows create data and tool-use risks that fit agentic AI governance concerns.
GDPRArt.32Sensitive personal data in prompts raises security and processing obligations under GDPR.

Assign clear governance for AI data use, approval, and accountability across the model and tool stack.


Key terms

  • Data Lineage: The record of how data moves across systems, applications, and workflows. In security operations, lineage shows where sensitive data propagates, which identities touch it, and how a compromise could spread across connected environments.
  • Prompt lineage control: Prompt lineage control is the enforcement of policy on AI prompts based on the origin and classification of the source data. It extends DLP beyond text scanning by connecting the prompt to the identity, source document, and destination tool, so policy can follow the data rather than the channel.
  • Shadow AI: AI agents, copilots, or connected tools operating without full visibility or governance from security teams. Shadow AI becomes an identity problem when those systems authenticate with unmanaged tokens, service accounts, or OAuth apps that can reach production resources.
  • Agentic workflow: An agentic workflow is a sequence of tasks executed by an AI agent with some level of tool access and decision authority. In security terms, the workflow matters because it can span multiple systems, identities, and permissions, which makes attribution and revocation harder than with ordinary automation.

What's in the full article

Cyberhaven's full post covers the operational detail this post intentionally leaves for the source:

  • Endpoint workflow examples showing how copy, paste, and browser submission are observed in practice.
  • Policy tuning guidance for allowing general AI use while restricting sensitive data classes.
  • Inventory and classification detail for AI tools that train on submitted content or retain conversation history.
  • Linea AI workflow coverage for agentic pipelines and downstream output monitoring.

👉 Cyberhaven's full post covers lineage-based enforcement, AI tool inventory, and agentic workflow monitoring.

Deepen your knowledge

NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, secrets management, and agentic AI identity. It gives security and identity practitioners a shared foundation for governing modern access pathways and non-human workflows.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 2, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org