By NHI Mgmt Group Editorial TeamBased on WitnessAI: “The Coming AI Architecture Shake-Up: What Enterprises Must Prepare For in 2026” (December 17, 2025)

TL;DR: Enterprises are moving from copilot-style add-ons toward AI-first architectures where agents orchestrate tools through protocols like MCP, while GPU provisioning delays and static capacity assumptions expose a growing mismatch between AI adoption and infrastructure reality, according to WitnessAI. The governance problem is no longer just adding AI features, but deciding how identities, access, and runtime control work when AI becomes the orchestration layer.


At a glance

What this is: This is WitnessAI's view that AI-first architectures will replace copilot-style add-ons, with MCP becoming a standard interface for agent access and infrastructure limits forcing a reset in AI deployment assumptions.

Why it matters: It matters because IAM, PAM, and NHI teams will have to govern agent-to-tool access, runtime privilege, and service capacity as part of the identity model rather than treating AI as a feature layer.


Context

AI-first architecture changes the identity problem because the system that decides what to do is no longer the same system that stores data or exposes functions. In practice, the control question shifts from app login to runtime authorisation across tools, data stores, and orchestration layers.

MCP is the protocol pattern that makes this shift concrete. Instead of custom integrations for each application, AI agents can call standardised tool interfaces, which means governance has to cover who can expose tools, what data those tools can return, and how those permissions are constrained.

The operational gap is not just architectural elegance. It is the difference between a copilot model that stays inside an application boundary and an AI-first model where the agent becomes the interface to enterprise systems, with all the access and lifecycle consequences that follow.


Key questions

Q: How should security teams govern MCP servers used by AI coding assistants?

A: Treat MCP servers as privileged trust boundaries, not simple data sources. Security teams should classify each server by the authority it can influence, sanitize any user-generated or third-party content before delivery, and limit the agent’s tool access so malicious context cannot easily become destructive action.

Q: Why do AI-first architectures change identity governance for enterprise systems?

A: Because the agent, not the application, becomes the entity that chooses which systems to query and which actions to take. Traditional governance assumes access is tied to one app or one workflow. In AI-first designs, the effective privilege set is distributed across tools, orchestration, and data access, so governance must follow the runtime path.

Q: What breaks when AI systems need multiple enterprise tools at runtime?

A: Static access models break first, because the system cannot predict every tool combination in advance. That creates pressure to overprovision, especially when the agent must complete a task across databases and applications. The result is broader-than-intended access unless teams enforce task-scoped control and separate duties across the tool chain.

Q: How do infrastructure constraints affect AI identity controls?

A: When GPU capacity is slow to provision or fixed in advance, teams are tempted to relax security controls to keep services online. That can turn temporary access exceptions into standing privilege. The governance question is therefore not just whether AI can run, but whether the control model survives under load.


Technical breakdown

Why the copilot model breaks under AI-first orchestration

The copilot model assumes AI is a feature attached to an application, with the application still acting as the primary control point. AI-first architecture reverses that relationship: the AI system selects tools, sequences actions, and pulls data across multiple systems to complete a task. That means identity is no longer bound to one front end or one workflow. Governance must account for the agent as the orchestrator, not just the app as the protected resource. In NHI terms, the control boundary moves from the user-facing session to the machine identity and tool permissions that the agent uses behind the scenes.

Practical implication: map every AI workflow to the underlying tool and data permissions it can invoke, not just the application where the prompt begins.

What MCP servers change for enterprise access control

MCP creates a standard interface for AI systems to access enterprise services, much like APIs standardised application-to-application communication. The governance consequence is that the MCP server becomes an identity enforcement point, not just a transport layer. Teams have to decide which tools are published, whether access is scoped per agent or per task, and how responses are filtered before they reach the model or downstream agent logic. Without those controls, MCP reduces integration friction while increasing the blast radius of overbroad tool exposure.

Practical implication: treat MCP server design as part of access governance and inventory every tool the server exposes to agents.

Why GPU bottlenecks become an identity governance issue

The article's infrastructure point is that AI workloads do not behave like conventional cloud workloads. GPU capacity may take 20 to 30 minutes to provision and is often allocated statically, which breaks assumptions about elastic scaling. For identity teams, that matters because access policy, service availability, and approval workflow timing all depend on predictable runtime behaviour. When AI systems stall or fail under load, teams often respond by widening access, bypassing gating, or hard-coding entitlements to keep services running. That is a governance failure, not just an infrastructure delay.

Practical implication: align AI access models with real capacity constraints so emergency workarounds do not turn into standing privilege.


Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

AI-first architecture turns the agent into the new control plane: once AI systems orchestrate enterprise tools, identity governance can no longer stop at the application boundary. The real subject is not whether AI is present, but where authorisation is enforced when the agent chooses the sequence of actions. That makes tool access, runtime scoping, and delegation policy the core governance objects for IAM, PAM, and NHI teams.

MCP servers create a standardised identity surface, not just an integration layer: standardisation helps adoption, but it also concentrates risk because the same interface can expose many enterprise systems to many agents. The important question is not whether MCP exists, but whether the server design encodes tool-level boundaries, data minimisation, and revocation paths. Practitioners should treat MCP inventory as part of identity governance, not as a separate engineering concern.

Least privilege was designed for known workloads, not for agents that assemble tasks at runtime: that assumption fails when an AI system decides which tools to use only after receiving context, because the effective privilege set is not stable at provisioning time. The implication is that access governance has to move from static entitlement thinking to task-scoped control of the agent's tool chain.

AI capacity constraints now shape governance outcomes: GPU bottlenecks and static allocation models create pressure to weaken controls when adoption spikes. When the infrastructure cannot keep pace with demand, organisations often route around approval gates to preserve service continuity. That makes capacity planning part of identity risk management, not a separate infrastructure discussion.

Runtime tool trust debt: AI-first adoption accumulates hidden trust in the interfaces between agents, protocols, and enterprise systems. Every newly exposed tool expands the set of actions an AI system can initiate, and every shortcut taken to keep it working increases the amount of access that must later be re-reviewed. Practitioners should treat this as an identity governance debt problem, not an integration milestone.

From our research library:

What this signals

Runtime tool trust debt: AI-first architecture expands the number of enterprise actions that can be initiated through agentic systems, so identity governance has to move closer to the tool layer. The practical change for programmes is that inventory, scoping, and revocation now matter as much for agent access paths as they do for human accounts.

Identity teams should expect MCP-style interfaces to become a governance choke point. If the organisation cannot answer which tools an agent can invoke, which data each tool returns, and who owns revocation, the architecture is already ahead of the control model.

The capacity story and the identity story are now linked. When AI services stall because GPU resources are not available on demand, the operational pressure often lands on security teams as a request to widen access or bypass approval steps, which turns infrastructure strain into entitlement risk.


For practitioners

  • Define agent-scoped access policies Classify each AI workflow by the exact tools, datasets, and actions it may invoke, then separate those permissions by task or agent role instead of reusing broad application accounts.
  • Inventory MCP-exposed tools and data paths Document every service published through MCP, including read and write capabilities, downstream data returned, and the business owner responsible for revocation.
  • Replace app-centric review logic with runtime governance Review whether your access reviews and entitlement approvals still make sense when AI systems orchestrate actions across multiple systems in a single session.
  • Align capacity planning with security controls Test how your AI services behave under GPU scarcity and demand spikes so that teams do not respond by granting standing access or bypassing approvals to preserve uptime.

Key takeaways

  • AI-first architecture shifts governance from application-centric controls to runtime control of agent tool use across enterprise systems.
  • MCP servers make agent access standardised, which improves integration but also creates a new identity enforcement point that must be inventoried and governed.
  • Infrastructure constraints such as slow GPU provisioning can push teams toward risky access exceptions, so capacity planning and identity governance now have to be managed together.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF, NIST Zero Trust (SP 800-207) and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseAI agents orchestrating tools create privilege and delegation risks central to this article.
Recommendation — Scope agent permissions to the minimum tool set and separate orchestration privilege from data access.
OWASP Non-Human Identity Top 10NHI-05 — Overprivileged NHIAgentic systems rely on machine credentials and tool access that can easily exceed intended scope.
Recommendation — Audit machine and agent credentials for excess tool reach and remove standing access paths.
NIST AI RMFGOVERN — AI Governance and AccountabilityThe article is about the governance layer that must sit above AI-first orchestration choices.
Recommendation — Assign governance owners for AI orchestration, tool exposure, and approval boundaries.
NIST Zero Trust (SP 800-207)Resource access and segmentationAI-first tool access needs continuous verification and segmented exposure across systems.
Recommendation — Segment AI tool access and enforce continuous verification at each decision point.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe core issue is how access permissions are assigned and controlled for agent-driven workflows.
Recommendation — Review entitlements for agent workflows and remove any permissions not required at runtime.

Key terms

  • AI-first Architecture: AI-first architecture is a design pattern where the AI system orchestrates work across tools instead of sitting beside an application as a feature. For identity teams, that changes the governance target from a single app session to a runtime chain of actions, data calls, and delegated access.
  • MCP Server: An MCP server is a tool endpoint that connects an AI agent to external systems and data sources through Model Context Protocol. Because it extends what the agent can reach, it becomes part of the identity and access surface and must be reviewed like any other privileged connector.
  • Agent-scoped access: Agent-scoped access limits what an AI system can do based on the specific task or role it is performing. Unlike user-centric access models, it has to account for dynamic tool selection at runtime, which makes entitlement design and revocation more complex.
  • Runtime Governance: Runtime governance is the set of controls that verify what a system or agent is actually doing after deployment. It combines monitoring, authorization checks, and access validation so teams can detect drift, misuse, or excessive privilege in motion rather than assuming build-time policy still holds.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org