TL;DR: Verizon’s 2026 Data Breach Investigations Report, published by Sentra, finds Shadow AI is now the third most common insider DLP event, third-party breaches rose 60% to 48% of all breaches, and vulnerability exploitation drove 31% of incidents, showing that data exposure is increasingly governed by visibility gaps rather than isolated technical failures. The decisive control problem is knowing what sensitive data exists, where it lives, and who can reach it before AI, vendors, or attackers do.
At a glance
What this is: This analysis argues that the 2026 DBIR points to a single governance gap: organisations cannot answer basic questions about sensitive data visibility and access fast enough to contain modern exposure paths.
Why it matters: That matters to IAM practitioners because data access, third-party reach, and AI usage now depend on the same upstream controls that determine whether identities, workloads, and users can move sensitive data out of scope.
By the numbers:
- Verizon’s 2026 DBIR covers more than 22,000 confirmed breaches and 31,000 total incidents in the twelve months ending October 2025.
- Third-party breaches now account for 48% of total breaches, up 60% year over year.
- Vulnerability exploitation accounted for 31% of data breaches in the study period.
👉 Read Sentra’s analysis of the 2026 DBIR through an AI data readiness lens
Context
AI data readiness is the ability to know what sensitive data exists, where it lives, and who or what can reach it before that data is moved, shared, or exposed. In this 2026 DBIR analysis, the issue is not a single control failure but a governance gap that spans user behaviour, third-party access, and vulnerability response.
For IAM, PAM, and data security teams, that gap matters because modern exposure often starts with access that is technically valid but operationally unsafe. When employees, vendors, or attackers can reach sensitive information without accurate classification and entitlement context, the control plane is already behind the risk.
The patterns in the DBIR are not atypical for enterprises that adopted AI, cloud collaboration, and third-party connectivity faster than they matured their data governance model. They are increasingly common where identity and data controls were designed separately rather than as one access problem.
Key questions
Q: How should security teams govern employee use of public AI tools in the browser?
A: They should treat browser AI use as an identity and data-control problem, not just an acceptable-use issue. The team needs visibility into what was pasted, which account was active, whether the content was sensitive, and whether policy enforcement occurred before the data left the organisation. Controls that only inspect network events will miss the real decision point.
Q: Why do third-party breaches remain so difficult to reduce?
A: They remain difficult because many programmes assess vendor security posture without continuously validating what data the vendor can actually reach. A clean questionnaire does not remove standing access, stale API credentials, or forgotten integrations. The control failure is lifecycle management of delegated access, not simply weak due diligence.
Q: How can organisations prioritise vulnerabilities using data context?
A: They should score exposures by the sensitivity of the data protected by the affected asset, not by technical severity alone. A misconfiguration that touches regulated customer data deserves faster action than one on an empty test system. This shifts remediation from volume-based triage to exposure-based decision making.
Q: Who is accountable when sensitive data leaves through a vendor, API, or misconfigured system?
A: Accountability usually sits with the business owner of the data, the identity or platform team that granted access, and the vendor manager if external trust was involved. Frameworks such as Zero Trust and least privilege make that shared responsibility harder to ignore because they require continuous verification of access, not one-time approval.
Technical breakdown
Shadow AI as an identity and data leakage path
Shadow AI is not only a data loss issue. It is an access and classification issue that begins when employees use external AI services with data they are already entitled to see but not entitled to export. Standard DLP assumes a defined boundary such as email, web upload, or removable media, yet AI-assisted work often looks like ordinary browser traffic to a legitimate service. That makes the policy question upstream: was the data classified, was access right-sized, and were the approved AI endpoints governed before usage spread?
Practical implication: classify sensitive data and bind AI access policy to entitlement context before relying on DLP to catch leakage.
Third-party exposure and the reach of delegated access
Third-party risk is often treated as a due diligence exercise, but the breach path is usually delegated access plus poor lifecycle control. A vendor can pass a questionnaire and still retain a live connection to sensitive records long after the business need changed. In identity terms, the issue is persistent reach without continuous review, especially when cloud accounts, APIs, and shared services are involved. The problem is not whether the vendor has controls in the abstract. It is whether their access to your data remains justified every day it stays open.
Practical implication: inventory third-party entitlements against the data they can actually reach and review them as continuously as human access.
Why vulnerability severity depends on data context
Vulnerability exploitation becomes more damaging when security teams cannot tell what data sits behind the exposed system. A misconfigured bucket, unpatched service, or overpermissioned workload is not equally severe in every case. The difference is the sensitivity and reach of the data attached to it. Without that context, remediation queues flatten into noisy lists and patch timing becomes detached from business impact. The DBIR finding is less about patching speed alone and more about prioritisation failure caused by poor data visibility.
Practical implication: enrich vulnerability and configuration findings with data sensitivity and identity reach so remediation is driven by exposure, not volume.
Threat narrative
Attacker objective: The objective is to obtain or move sensitive business data from environments where access was valid but governance was incomplete.
- Entry occurs through normalised AI usage, third-party connectivity, or exposed infrastructure rather than a single dramatic compromise.
- Escalation happens when users, vendors, or attackers can reach sensitive data that was never tightly classified or access-scoped.
- Impact follows when data leaves the organisation through external AI services, third-party paths, or exploited resources before governance catches up.
NHI Mgmt Group analysis
AI data readiness is now an identity governance problem, not just a data governance problem. The DBIR’s Shadow AI finding shows that employees are already moving sensitive data through sanctioned identities and unsanctioned services. That means the access problem starts with what a user can reach, not only with what they later leak. IAM, IGA, and data governance must be treated as one operating model, because classification without entitlement control still leaves data reachable. Practitioners should treat AI usage as an entitlement decision, not a content-filtering problem.
Persistent third-party reach is the hidden governance gap behind much of the breach growth. A vendor questionnaire can describe controls, but it cannot prove that delegated access is still necessary or scoped tightly enough. The real failure mode is standing access to sensitive data that outlives the business relationship or the project need. This is especially relevant where vendor access is mediated through cloud identities, service accounts, or API credentials. Practitioners should continuously validate third-party data reach rather than trusting point-in-time assurances.
Visibility into data context is becoming a prerequisite for meaningful security prioritisation. The DBIR’s vulnerability finding is important because it shows how patching queues become less useful when teams cannot tell which assets protect sensitive records. That is a control-design problem, not a workflow problem. If the organisation cannot connect exposures to the data they guard, it cannot rank risk accurately. Practitioners should make data sensitivity part of every remediation decision, not an afterthought.
Overexposed data is the new blast-radius multiplier. The named concept here is the condition where access breadth, third-party connectivity, and AI usage combine to turn routine entitlement into enterprise-scale exposure. This is not a single-tool failure. It is the result of separate programmes making incompatible assumptions about who can reach what. The practical conclusion is that data reachability must be governed as tightly as authentication.
OWASP-NHI becomes relevant the moment AI tools and third parties are granted enduring access to sensitive datasets. Even in a data-first article, the identity layer is what determines whether access can be constrained, reviewed, and revoked before exposure occurs. That makes lifecycle controls, least privilege, and credential hygiene central to AI data readiness. Practitioners should align data governance with identity governance rather than treating them as separate disciplines.
What this signals
Overexposed Data becomes the controlling risk model: as AI adoption, vendor connectivity, and cloud exposure converge, the meaningful question is no longer whether a control exists, but whether it is tied to the data an identity can actually reach. Teams that cannot join entitlement, classification, and sensitivity data will keep reacting after exposure instead of preventing it.
For identity and security programmes, the signal is clear. AI governance, third-party access review, and vulnerability management will increasingly converge around data reachability, which means entitlement inventories and classification coverage become operational prerequisites rather than audit artefacts. The organisations that can evidence control over reach will be able to prioritise faster and defend decisions more credibly.
For practitioners
- Classify sensitive data before AI access expands Map the datasets employees can reach, then classify the records that should never be exposed to external AI systems or unmanaged browser-based tools. Tie that classification to policy enforcement so sensitive data is blocked upstream, not only detected after submission.
- Review third-party entitlements against actual data reach List every vendor, API, and shared service that can touch sensitive records, then verify whether the access is still required and properly scoped. Remove standing access where the business need has ended and recertify delegated access on a continuous cadence.
- Enrich remediation with data sensitivity context Add sensitivity labels, ownership, and identity reach to vulnerability and misconfiguration queues so teams can rank exposures by business impact. Use that context to separate empty or low-risk assets from systems that protect regulated or operationally critical data.
- Govern AI usage as an entitlement pathway Treat approved AI tools, browser extensions, and shadow services as access destinations that need policy, logging, and review. Align IAM and DLP rules so employee identities cannot move classified content into external services without explicit controls.
Key takeaways
- The DBIR findings point to one recurring failure: organisations do not know enough about their sensitive data to govern how AI, vendors, and attackers can reach it.
- Shadow AI, third-party breaches, and vulnerability exploitation all become more damaging when access decisions are made without current data context and entitlement visibility.
- Practitioners should connect IAM, data classification, and remediation workflows so sensitive data exposure is controlled before it becomes an incident.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Non-Human Identity Top 10 address the attack surface, NIST CSF 2.0, NIST SP 800-53 Rev 5 and NIST Zero Trust (SP 800-207) set the technical controls, and ISO/IEC 27001:2022 define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | PR.DS-1 | Data classification and protection are central to the AI data readiness gap described here. |
| NIST SP 800-53 Rev 5 | AC-6 | Least privilege is directly relevant to overexposed data and third-party access. |
| OWASP Non-Human Identity Top 10 | NHI-03 | The article’s access and lifecycle issues intersect with NHI governance for service accounts and API paths. |
| NIST Zero Trust (SP 800-207) | 3.1 | Continuous verification fits the need to reassess who can reach data across AI and vendor pathways. |
| ISO/IEC 27001:2022 | A.5.9 | Inventory and ownership of information assets underpin the data visibility gap discussed in the article. |
Maintain an accurate information asset inventory so access decisions are based on current data ownership and sensitivity.
Key terms
- AI Data Readiness: AI Data Readiness describes whether an organisation can safely expose data to AI systems without losing control over sensitivity, purpose, or access scope. It combines discovery, permission management, and continuous oversight so data use remains aligned to governance expectations.
- Shadow AI: AI agents, copilots, or connected tools operating without full visibility or governance from security teams. Shadow AI becomes an identity problem when those systems authenticate with unmanaged tokens, service accounts, or OAuth apps that can reach production resources.
- Delegated Access: Delegated access is permission granted to one identity to act on behalf of another user, service, or system. In NHI environments, this usually appears in OAuth-connected apps and automation tooling. It is powerful, but it must be tightly scoped and reviewed because it can persist long after the original business need ends.
- Dependency Reachability: Dependency reachability is the question of whether a vulnerable library or function can actually be invoked in the deployed application path. It matters because not every disclosed package flaw creates equal risk. Teams use it to separate theoretical exposure from issues that can be exploited in practice.
What's in the full article
Sentra's full article covers the operational detail this post intentionally leaves for the source:
- Sentra's breakdown of how AI data readiness maps to Shadow AI, third-party risk, and vulnerability prioritisation in practice.
- The article's explanation of how continuous data classification changes remediation timing and access decisions across cloud, SaaS, and on-prem environments.
- Practical examples of how overexposed data becomes a governance issue rather than a single DLP alert.
- The source's full framing of how AI data readiness fits into broader enterprise security operations.
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It is designed for practitioners who need to connect identity controls to wider security and data exposure decisions.
Published by the NHIMG editorial team on August 19, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org