TL;DR: AI data leaks now span prompts, coding assistants, training data, and autonomous agent workflows, and 20% of organizations with a breach said shadow AI was involved in 2025, according to WitnessAI. The real issue is not just leakage, but that existing DLP and CASB controls were never built for intent-driven AI interactions or machine-speed data movement.
At a glance
What this is: This guide explains why AI data leaks are a distinct enterprise risk and shows how shadow AI, model memory, coding assistants, and autonomous agents expand exposure beyond traditional DLP coverage.
Why it matters: IAM, security, and governance teams need to treat AI usage as an access and data-control problem because the same controls that govern static systems miss intent-driven interactions and machine-speed exfiltration.
By the numbers:
- 20% of organisations that suffered a data breach said the security incidents involved shadow AI in 2025.
Context
AI data leaks are an access and governance problem as much as a data protection problem. When employees, copilots, coding assistants, or autonomous agents move sensitive information into AI systems, the organisation loses visibility over where that data goes and who can recover it.
Traditional DLP, CASB, and SSE controls were built for labels, keywords, and perimeter inspection. AI interactions are contextual, probabilistic, and often happen in tools that security teams do not fully see, which is why shadow AI and sanctioned AI create the same core exposure when governance is missing.
Key questions
Q: What breaks when employees use unapproved AI tools with company data?
A: Governance breaks because the organisation loses visibility into where data and secrets are going, who can access them, and how they are being reused. Unapproved tools can copy credentials into unmanaged workflows, which weakens revocation and makes audit trails incomplete. The result is shadow access outside the main identity programme.
Q: Why do AI data leaks create a different risk than traditional data loss?
A: Traditional data loss controls assume static content, known channels, and predictable movement. AI changes all three by turning prompts, model responses, embeddings, and agent actions into data paths that can bypass keyword rules and perimeter monitoring. The risk is not only exfiltration, but exposure through a workflow that looks normal to the user.
Q: How can organisations tell whether AI governance is actually working?
A: Organisations can tell AI governance is working when they can inventory every agent, explain its purpose, show who owns it, and prove that permissions are tightly scoped. If those four things are missing, the programme has policy language but not operational control. Auditors will notice the gap quickly.
Q: What is the difference between sanctioned AI and shadow AI?
A: Sanctioned AI has gone through procurement, legal, and security review, with defined ownership and policy controls. Shadow AI bypasses those gates, often using existing browser sessions or SaaS permissions to process sensitive data outside approved oversight. The difference is not the model. It is the control path.
Technical breakdown
Why DLP and CASB miss AI data leaks
AI data leaks differ from classic exfiltration because the data move is often hidden inside a conversational or workflow context rather than a file transfer. DLP depends on stable labels and recognizable patterns, while AI prompts, responses, embeddings, and tool calls can transform the same content into forms those tools cannot inspect. That creates a control gap at the point of interaction, not just at rest or on the network edge. Traditional controls also assume deterministic behaviour, but LLMs and agents can produce different outputs from similar inputs, which undermines signature-based detection.
Practical implication: shift detection and control to the interaction layer, not only the perimeter or storage tier.
How shadow AI and coding assistants expand the exposure surface
Shadow AI is unmanaged AI use that bypasses approved workflows, but the risk is not limited to unsanctioned tools. Coding assistants and embedded copilots regularly process source code, credentials, and confidential business content as part of ordinary work, which makes leakage an operational byproduct of productivity. The article also notes that developers paste hardcoded API keys and secrets into assistants, creating immediate operational exposure. In identity terms, these tools blur the line between user action and machine-assisted disclosure, so governance has to cover both sanctioned and unsanctioned use.
Practical implication: classify and govern AI access by use case and data sensitivity, not by whether the tool is officially approved.
Why agentic workflows change the trust model
Autonomous agents are different from user-driven AI because they can query systems, call APIs, and combine outputs across multiple tools without a human deciding each step. That turns a single interaction into a chained data movement path. The article’s Model Context Protocol example shows another issue: if the server does not carry user context, it may grant similar access across users, which weakens least-privilege assumptions at runtime. This is not just more usage, it is a different trust model where access, context, and execution can drift inside one workflow.
Practical implication: treat agent tool access and context propagation as governance controls, not implementation details.
Threat narrative
Attacker objective: The objective is to obtain sensitive enterprise information through AI workflows without triggering the controls built for conventional data movement.
- Entry occurs when employees paste confidential data into public AI tools, coding assistants, or embedded copilots as part of normal work.
- Credential access or disclosure follows when prompts, outputs, or agent tool calls expose source code, secrets, customer records, or memorized training data.
- Escalation happens when autonomous agents use API and system access to combine data sources and move sensitive information faster than human review can intervene.
- Impact is leakage of regulated, proprietary, or operational data into systems the organisation does not control, with regulatory and brand consequences.
Breaches seen in the wild
- Vercel Context.ai OAuth Supply Chain Breach: Shadow AI app Context.ai OAuth integration exposes Vercel customer data via unmanaged third-party token.
Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.
NHI Mgmt Group analysis
AI data leaks expose a governance gap, not just a content filtering problem. The failure is not simply that employees share too much with tools. It is that most enterprise controls still assume sensitive data moves in predictable formats through known channels. When the interaction itself is conversational, probabilistic, and sometimes agent-driven, the control model has to govern intent, context, and execution, not just content.
Shadow AI and sanctioned AI are operationally different but governance-equivalent. The article shows the same risk pattern whether the usage is approved or unsanctioned: sensitive data moves into systems the organisation cannot fully inspect or constrain. That means policy boundaries matter, but visibility and enforcement matter more. The practitioner conclusion is that AI governance cannot stop at approval lists or acceptable-use language.
Model Context Protocol makes context loss a governance issue. The architectural problem is not only tool access, but the absence of stable user context across server interactions. When identical access is granted regardless of who initiated the request, least privilege becomes harder to express and harder to audit. The implication is that AI governance must account for context propagation, not just authentication at the edge.
Intent-based control is the right concept for AI data movement. Traditional classification assumes the risk is visible in the data object itself. AI workflows make the user’s purpose and the surrounding context just as important as the content. That shifts governance toward behavioural analysis, inline policy, and runtime inspection, which is where enterprise AI security is headed.
AI usage visibility is now a core identity-security requirement. If security teams cannot see prompts, tool calls, responses, and agent actions, they cannot govern the identities operating through those channels. For IAM, PAM, and NHI programmes, AI data leakage is another sign that access control is becoming runtime control, and the programme has to mature accordingly.
From our research library:
- Organisations that rely heavily on static credentials reported a 20-percentage-point increase in security incidents compared with those with low reliance, according to the 2026 Infrastructure Identity Survey.
- Only 5.7% of organisations have full visibility into their service accounts, according to the Ultimate Guide to NHIs.
- Read next: Shadow AI and AI Agent Discovery Guide
What this signals
AI data leak governance has to move from perimeter logic to interaction logic. Once prompts, responses, and tool calls become the primary movement path for sensitive information, controls based on file inspection and network choke points lose relevance. Security teams should expect this shift to reshape how they evaluate AI usage, audit trails, and data handling inside identity programmes.
Shadow AI is best understood as unmanaged access to an AI workflow, not just unsanctioned software. That framing matters because it aligns AI governance with the same programme questions used for NHI and privileged access: who can act, under what context, with what data, and with what evidence afterward. The organisation that can answer those questions can govern AI; the one that cannot is still guessing.
For practitioners
- Establish an AI acceptable-use policy Define which AI tools, data classes, and use cases are permitted, then pair the policy with enforcement so it cannot be bypassed through personal accounts or shadow tools.
- Extend visibility across all AI surfaces Track browser use, desktop apps, coding assistants, embedded copilots, and agent connections so prompts and outputs are visible where data first enters or leaves the environment.
- Classify by intent and context Evaluate what the user is trying to do, not just the words in the prompt, so legitimate analysis does not trigger the same controls as pre-disclosure sharing or exfiltration.
- Tokenize sensitive data before AI submission Replace high-risk values such as API keys, customer records, and regulated identifiers before they reach third-party models, then rehydrate only when policy allows.
- Govern agent tool access as privileged access Treat autonomous agent workflows as high-risk access paths, with scoped permissions, pre-execution checks, and audit trails for both tool requests and outputs.
Key takeaways
- AI data leaks are now a governance issue because ordinary work interactions can move sensitive information into AI systems outside traditional visibility and control.
- WitnessAI cites 20% of breached organisations in 2025 as having incidents that involved shadow AI, which shows the exposure is already material.
- The practical response is to govern AI usage by intent, context, and runtime policy rather than relying on static data labels and perimeter tools.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Non-Human Identity Top 10 | NHI-02 — Secret Leakage | The article repeatedly cites AI workflows exposing secrets, code, and confidential data. |
| NHI-10 — Human Use of NHI | Employees using AI tools as productivity channels create the misuse pattern discussed here. | |
| Recommendation — Scan AI workflows for secret leakage and block sensitive values before they reach external models. Restrict human-driven AI use paths that bypass approved data-handling and audit controls. | ||
| OWASP Agentic AI Top 10 | ASI02 — Tool Misuse | Autonomous agents in the article call APIs and access systems beyond simple chat use. |
| Recommendation — Constrain agent tool access so runtime actions cannot exceed the approved workflow. | ||
| NIST CSF 2.0 | PR.AA-05 — Access Permissions, Entitlements and Authorizations | The governance gap is fundamentally about who can move data through AI-enabled access paths. |
| Recommendation — Review AI entitlements and tighten access permissions for prompts, tools, and downstream data flows. | ||
| MITRE ATT&CK | TA0006;TA0010 — Credential Access; Exfiltration | The article describes credentials and sensitive data leaving through AI-assisted workflows. |
| Recommendation — Map AI leakage paths to credential access and exfiltration stages to prioritise monitoring and containment. | ||
Key terms
- AI data leakage: AI data leakage occurs when sensitive business information is exposed through prompts, outputs, or copied content in AI-assisted workflows. In browser-driven work, the risk is often accidental rather than malicious, so governance depends on data rules, usage policy, and session controls.
- Shadow AI: AI agents, copilots, or connected tools operating without full visibility or governance from security teams. Shadow AI becomes an identity problem when those systems authenticate with unmanaged tokens, service accounts, or OAuth apps that can reach production resources.
- Intent-based classification: Intent-based classification evaluates what a user or system is trying to do, not just what text or file is present. In AI governance, it distinguishes routine work from risky interaction by reading context, purpose, and sensitivity. That matters when regulated data is handled conversationally rather than through formal file transfer.
- Agentic workflow: An agentic workflow is a sequence of tasks executed by an AI agent with some level of tool access and decision authority. In security terms, the workflow matters because it can span multiple systems, identities, and permissions, which makes attribution and revocation harder than with ordinary automation.
Deepen your knowledge
NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org