By NHI Mgmt Group Editorial TeamBased on WorkOS: “Why most enterprise AI projects fail — and the patterns that actually work” (July 22, 2025)

TL;DR: A S&P Global survey of more than 1,000 enterprises found 42% abandoned most AI initiatives in 2025, while the average organisation scrapped 46% of AI proofs of concept before production, pointing to cost, privacy, and security failures, according to WorkOS and S&P Global. The real constraint is not model quality alone, but whether governance, data readiness, and human operating models can survive production pressure.


At a glance

What this is: This is an analysis of why enterprise AI initiatives stall, with the key finding that governance, data plumbing, operating model design and product discipline matter more than model sophistication.

Why it matters: It matters because IAM, NHI and security teams are increasingly being asked to support AI systems that fail in the handoff between pilot and production, where authentication, oversight, access control and workflow ownership become decisive.

By the numbers:

  • 42% of companies abandoned most of their AI initiatives in 2025, up from just 17% in 2024, according to WorkOS.
  • The average organization scrapped 46% of AI proof-of-concepts before they reached production, according to WorkOS.
  • Over 80% of AI projects fail, according to WorkOS.
  • Lumen Technologies projects $50 million in annual savings from AI tools, according to WorkOS.

Context

Enterprise AI programmes often fail at the point where a pilot meets production reality. The article argues that the problem is not model quality alone, but the lack of a governed path through authentication, compliance, data readiness and operational ownership.

For IAM and identity security teams, this is a lifecycle problem as much as a technology problem. AI initiatives can look functional in a sandbox while still lacking the access controls, human oversight and workflow integration needed for safe, durable deployment.


Key questions

Q: Why do AI ROI models often fail after a successful pilot?

A: Pilots usually measure activity, not production durability. Once AI moves into governed environments, data controls, access oversight, compliance logging and operating costs change the return profile. If those factors were not baselined early, the programme looks profitable in testing but underperforms in production.

Q: Should organisations prioritise AI data governance before scaling AI adoption?

A: Yes. Organisations that scale AI before establishing discovery, classification, monitoring, and policy enforcement are effectively expanding the attack surface faster than they can govern it. AI adoption should be matched with controls that follow the data lifecycle, otherwise compliance, exposure, and misuse risks compound as usage grows.

Q: What do organisations get wrong about human oversight in agentic AI?

A: They confuse a named reviewer with effective oversight. Real oversight requires training, escalation practice, and decision authority under pressure. If approvers have never rehearsed the scenario, they are likely to trust the system too quickly or miss the moment when denial is the safer outcome.

Q: How do security teams decide whether an AI workload is ready for production?

A: Use a governance test, not a marketing test. The workload is ready only if its models, dependencies, data sources, runtime controls, and resource limits are known, approved, and continuously monitored. If any of those elements are opaque, the deployment is still experimental from a security perspective.


Technical breakdown

Why AI pilots stall at production handoff

Many enterprise AI initiatives start in controlled environments that hide the real integration burden. The model may function in isolation, but production use demands secure authentication, compliance workflows, user training, and ownership across infrastructure, data and operations. That is where projects often slow or stop. The article’s pattern is not model failure, but deployment failure: the technical prototype is ready before the organisation is ready to absorb it. In identity terms, the programme has not defined who can trigger, approve, observe and override the AI system once it leaves the lab.

Practical implication: treat production readiness as an identity and workflow problem, not just a modelling milestone.

Why data plumbing determines whether AI can be governed

Enterprise AI depends on trusted data inputs, and the article makes clear that poor data quality, weak metadata, retention gaps and fragmented pipelines undermine outcomes before the model ever reaches users. Retrieval-augmented generation and similar systems are especially sensitive because they operationalise data quality at runtime, not just in training. If data is incomplete, stale or poorly governed, the AI system will surface bad answers, create trust issues and force rework. The governance burden therefore shifts from one-time preparation to continuous control over the data estate feeding the AI stack.

Practical implication: establish data governance, lineage and retention controls before scaling AI into customer- or employee-facing workflows.

Human-AI collaboration is the control plane, not a fallback

The article shows that durable deployments are designed around explicit handoffs between people and systems. Human oversight is not an emergency brake added after launch; it is part of the operating model. That matters because enterprise workflows rarely tolerate full automation without loss of trust, quality or accountability. When AI drafts, summarises or recommends, humans still need defined approval points, override paths and feedback loops. In practice, the control plane is the collaboration design itself: who owns the decision, who reviews the output, and when the system may act independently.

Practical implication: define handoff points and approval boundaries before deployment, especially where AI output affects customer, financial or compliance decisions.


Threat narrative

Attacker objective: The failure condition is not a compromise by an external attacker, but the collapse of the programme’s path to production, leaving value unrealised and governance fragmented.

  1. Entry occurs when teams deploy AI pilots in isolated sandboxes that do not yet have secure authentication, compliance workflows or production-grade ownership.
  2. Credential and access issues emerge when the system must cross from prototype into shared enterprise infrastructure, where permissions, oversight and integration become necessary.
  3. Escalation happens as disconnected teams, shadow IT and ungoverned data pipelines multiply operational risk and make production deployment harder to control.
  4. Impact is stalled adoption, abandoned proofs of concept, duplicated tooling and wasted investment as the AI programme fails to move into durable use.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Governance drift, not model drift, is the real enterprise AI failure mode. The article shows that most programmes do not fail because the model is incapable, but because the organisation cannot govern its movement into production. Authentication, workflow ownership and compliance design lag behind experimentation, so the programme loses control at the moment it needs it most. For practitioners, the key question is whether the AI initiative has a production governance model, not just a technical demo.

Data readiness is now an identity and access problem as much as a data problem. When AI depends on governed data sources, the quality of access, lineage and retention controls becomes part of the system’s trust model. This is where identity security teams matter, because the AI stack can only be as reliable as the permissions and controls around the data it consumes. The implication is that data governance and access governance must be designed together.

Human oversight is not optional once AI is embedded in enterprise workflows. The article’s strongest signal is that successful organisations design explicit collaboration patterns between humans and machines before launch. That is the difference between accountable automation and unmanaged output. Human-AI operating model drift: the process changes faster than the governance model, and that mismatch is what turns a working prototype into a failed deployment. Practitioners should treat handoffs, approvals and overrides as core controls, not process decoration.

The enterprise AI market is moving toward operational discipline, not model theatre. The winning pattern is not who can produce the most impressive prototype, but who can sustain trust, ownership and measurable value after launch. That shifts evaluation criteria toward governance maturity, cross-functional coordination and observability. For IAM and security leaders, this validates a programme-level view: AI adoption is now a control environment issue, not just an innovation initiative.

Shadow AI and duplicate stacks are a governance smell, not a side effect. The article’s description of orphaned GPU clusters, duplicate vector databases and parallel MLOps efforts shows how quickly enthusiasm creates unmanaged identity and infrastructure sprawl. That is a warning sign that enterprise AI is escaping central oversight before value is proven. Practitioners should read uncontrolled AI sprawl as a signal that governance has not kept pace with adoption.

From our research library:

What this signals

Enterprise AI adoption now depends on governance depth, not experimentation volume. The organisations most likely to progress are the ones that can align data controls, operating model design and accountability before the model is exposed to users. That shifts the centre of gravity from model selection to programme discipline.

Human-AI operating model drift: when deployment speed outpaces approval design, the organisation gets a system that is technically functional but operationally fragile. The practical test is whether a team can explain who owns the AI output, who can stop it, and who is responsible when it changes business outcomes.


For practitioners

  • Define the production path before the pilot Map the route from proof of concept to live service, including authentication, compliance approvals, user training and support ownership. If the path is unclear, the pilot is not ready for scale.
  • Prioritise data readiness over model tuning Assign budget and timeline to data extraction, normalisation, metadata, quality monitoring and retention controls before expanding the model footprint.
  • Design human handoffs into the workflow Specify where humans approve, override or review AI output, and make those checkpoints visible in the operating model rather than implicit in the process.
  • Instrument AI as a managed product Track event logs, model output quality, feature null rates and user feedback through existing operational dashboards so drift becomes routine maintenance instead of a surprise.

Key takeaways

  • Enterprise AI initiatives fail most often at the point where the pilot meets production, not at the point where the model is first trained.
  • The article ties abandonment to cost, privacy and security pressure, with 42% of companies walking away from most AI initiatives in 2025 and 46% of proof-of-concepts scrapped before production.
  • The clearest path to durable AI adoption is stronger governance over data, workflow ownership and human oversight, not more emphasis on model tuning alone.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and CSA MAESTRO address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST AI RMFGOVERN — AI Governance and AccountabilityThe article is fundamentally about AI programme governance, ownership and accountability.
MANAGE — AI Risk ManagementThe article emphasises operational controls, monitoring and lifecycle management after deployment.
Recommendation — Establish AI governance roles, decision rights and escalation paths before scaling a model into production. Operationalise monitoring, ownership and drift response for AI services as part of the managed lifecycle.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsProduction AI depends on governed access to data, workflows and supporting systems.
Recommendation — Review access permissions and entitlements for AI data pipelines and service integrations before go-live.
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseAI systems can fail when delegated access and workflow permissions are not bounded in practice.
Recommendation — Constrain AI privileges to the minimum needed for the workflow and separate approval from execution.
CSA MAESTROGovernance and orchestrationThe article highlights the need to orchestrate AI workflows across people, data and controls.
Recommendation — Design orchestration so AI actions remain visible, reviewable and bounded by human governance.

Key terms

  • Runtime AI Governance: Runtime AI governance is control applied while the interaction is happening, rather than before deployment or after an incident. It combines discovery, policy enforcement, output inspection, and audit logging so that AI use can be managed in live enterprise conditions.
  • Hybrid AI Operating Model: An operating model that combines internal ownership of data and governance with external tooling or domain expertise. It is common where organisations want speed without losing control, but it only works when accountability for data quality, access, and decisions remains clearly assigned.
  • Data Readiness: Data readiness is the degree to which data is clean, governed, accessible for the right purpose, and traceable back to a known source. For AI programmes, it covers lineage, retention, quality, and access controls, because poor data quality becomes a governance failure at runtime.
  • Shadow AI: AI agents, copilots, or connected tools operating without full visibility or governance from security teams. Shadow AI becomes an identity problem when those systems authenticate with unmanaged tokens, service accounts, or OAuth apps that can reach production resources.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 8, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org