By NHI Mgmt Group Editorial TeamBased on WorkOS: “MCP-UI: A Technical Overview of Interactive Agent Interfaces” (September 8, 2025)

TL;DR: MCP-UI extends the Model Context Protocol with interactive web components, sandboxed iframe rendering, and event-based UI actions that let agents handle richer workflows directly in conversation, according to WorkOS. The governance issue is not prettier interfaces but a new identity boundary where tool permissions, event validation, and session trust need tighter control than text-only MCP patterns allowed.


At a glance

What this is: This is a technical overview of MCP-UI, an extension to MCP that moves agent interactions from text-only output to embedded, event-driven web interfaces with sandboxed rendering and structured UI actions.

Why it matters: It matters because interactive agent interfaces expand the identity and authorisation boundary, forcing IAM, NHI, and agent governance teams to treat UI events as controlled execution paths rather than passive presentation.


Context

MCP-UI changes the security problem from exchanging text with an agent to governing interactive actions that originate inside an embedded user interface. That shift matters because once a component can emit tool, intent, prompt, or link events, the trust boundary moves from model output to event validation and session control.

For identity teams, the core question is not whether the interface is richer. It is whether existing MCP governance patterns can safely handle UI-driven actions that look like user interaction but behave like delegated execution. That is a different governance problem from plain tool invocation.


Key questions

Q: What breaks when agent UIs can trigger actions directly?

A: When UI events can trigger actions directly, the agent loses its role as a policy gate and the interface becomes an uncontrolled execution path. That increases the risk of hidden privilege escalation, untraceable state changes, and confusing accountability. The control failure is not visual complexity, but bypass of the mediated decision point.

Q: Why do interactive agent interfaces increase authorisation risk?

A: They increase risk because a user-facing component can now initiate actions that previously required explicit tool invocation. That creates a larger trust boundary around the session and makes hidden privilege escalation easier if event handling is weak. The core issue is not the iframe itself, but the authorisation rules attached to the events it generates.

Q: How should teams validate embedded UI events in agent workflows?

A: Teams should validate sender origin, event schema, allowed action types, and the sensitivity of the downstream operation before accepting any UI event. The goal is to prevent a rendered component from becoming an ungoverned request generator. Validation should happen at the boundary where presentation becomes execution, not after the agent has already acted.

Q: What is the difference between sandboxed rendering and controlled event execution?

A: Sandboxed rendering isolates code execution from the host application, while controlled event execution governs what that isolated code is allowed to ask the agent or backend to do. A sandbox reduces direct compromise risk, but it does not authorise behaviour. Security teams need both isolation and policy enforcement, because one without the other leaves a gap.


Technical breakdown

UIResource and ui:// resources

MCP-UI extends MCP with a UIResource object that points to renderable content through a ui:// URI and a declared MIME type. That allows a server to send HTML, a remote URL, or a Remote DOM payload as a first-class resource. The protocol does not replace MCP's tool model. It adds a presentation layer that still depends on structured resource delivery, encoding, and client-side rendering rules. Practical implication: teams need to govern which resources can be surfaced to a client and how those resources are classified before they are rendered.

Practical implication: govern renderable resources as part of the agent trust boundary, not as inert UI assets.

Sandboxed iframe rendering and Remote DOM

MCP-UI supports three main rendering paths: inline HTML in a sandboxed iframe, external URL embedding through an iframe src, and Remote DOM for JavaScript-driven interfaces. Each path changes the security model. Inline HTML limits external dependencies. External iframes expand capability while increasing exposure to origin and content trust issues. Remote DOM reduces iframe overhead but requires tighter control of message flow and component behaviour. Practical implication: choose the rendering path based on the level of trust, interactivity, and origin isolation you can actually enforce.

Practical implication: set rendering policy by trust level, because each mode creates a different attack surface.

Event-based UI actions and structured intent

Instead of allowing embedded components to mutate state directly, MCP-UI uses structured events such as tool, intent, prompt, notify, and link. That preserves a boundary between presentation and execution, but it also means security now depends on event validation, origin checks, and downstream authorisation decisions. The event channel becomes the real policy enforcement point. If the action semantics are not tightly validated, a polished interface can become an execution proxy for untrusted input. Practical implication: treat every UI event as a policy decision, not a harmless click.

Practical implication: validate UI events as policy-bearing requests before they reach tool execution.


NHI Mgmt Group analysis

MCP-UI creates a new interface identity boundary, not just a better user experience. The meaningful change is that an agent can now carry richer interaction state through embedded components instead of plain text. That moves governance from message handling to event handling, because the interface itself can now initiate actions that look user-driven but behave like delegated execution. Practitioners should treat the UI layer as part of the identity plane, not a cosmetic front end.

Text-only MCP assumptions no longer hold once interfaces can emit structured intents. Traditional MCP governance assumes the agent response is the primary control surface. MCP-UI breaks that assumption by letting the interface generate tool calls, prompts, and links through sandboxed components. The implication is that least privilege, approval flow design, and session trust must now consider the UI as an active participant in authorisation, not a passive renderer.

Sandboxing reduces exposure, but it does not eliminate trust propagation. Iframes and Remote DOM keep code isolated from the host page, yet they still rely on event routing and origin validation to preserve control. That means the weakest point is often not the sandbox itself but the policy boundary around what events are accepted, transformed, or forwarded. Practitioners should focus on event trust, not just rendering isolation.

Interactive agent interfaces will force convergence between application security and identity governance. MCP-UI sits at the intersection of tool access, UI state, and agent workflow control, so the old separation between frontend behaviour and identity policy becomes harder to defend. This is where NHI governance starts to resemble application trust management, because the control question is no longer only who can call a tool, but which UI states are allowed to trigger that call. Practitioners should align interface design with explicit authorisation rules.

Ephemeral interaction states will become a governance problem before they become a usability advantage. The more an agent relies on interactive components, the more state is created and discarded inside a single conversational session. That state is harder to review after the fact and easier to mis-handle if validation is loose. The field will need stronger session-level controls for interactive agent workflows, especially where event semantics map to business actions. Practitioners should plan for governance at the moment of interaction, not after the conversation ends.

From our research library:

What this signals

Interactive UI turns MCP into a policy-bearing channel. The important governance shift is that the conversation surface now contains executable intent, not just language. That means agent programmes need controls for event provenance, session trust, and downstream action scoping before interactive components are allowed into production.

MCP-UI should be treated as an identity design problem first and a front-end problem second. Once a component can trigger a tool call, the trust model depends on who or what generated the event and whether that event is still valid at the moment of execution. Teams should align authorisation logic with UI state, not with presentation alone.

24,008 unique secrets were exposed in MCP configuration files in 2025 alone, the protocol's first year of widespread adoption. That scale shows how quickly protocol extensions can accumulate governance debt if teams focus on capability before boundary control.


For practitioners

  • Define a UI event policy boundary Classify which UI actions may emit tool, intent, prompt, notify, or link events, and require explicit validation before any event reaches downstream execution.
  • Restrict rendering modes by trust level Permit inline HTML, external iframe embedding, or Remote DOM only where the origin, content source, and execution model match the sensitivity of the workflow.
  • Add origin and message validation to every component Verify sender identity, expected message schema, and allowed action types before accepting postMessage or similar cross-context communication from embedded UI.
  • Map agent UI flows to approval requirements Identify which interactive workflows can trigger privileged tool calls and require step-up authorisation when a component can initiate business-impacting actions.
  • Separate presentation state from execution state Keep transient interface state from becoming implicit authorisation context, so a rendered component cannot silently carry trust into a tool invocation.

Key takeaways

  • MCP-UI extends the agent interface from text exchange to event-driven interaction, which changes the trust boundary that IAM and NHI teams need to govern.
  • Sandboxing helps contain rendering risk, but structured UI events still need validation because they can become execution requests.
  • The practical control question is no longer only what an agent can call, but which interface states are allowed to trigger those calls.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10, OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack and risk surface, while NIST CSF 2.0 and NIST Zero Trust (SP 800-207) set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-10 — Human Use of NHIMCP-UI lets human-facing interface events trigger non-human actions through the agent layer.
Recommendation — Limit user-triggered UI events that can flow into NHI execution and require validation before they reach tools.
OWASP Agentic AI Top 10ASI02 — Tool MisuseInteractive components can steer agents into unintended tool calls through structured events.
Recommendation — Constrain which interface events may invoke tools and block unapproved action paths.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsUI-triggered actions need explicit authorisation boundaries before they become execution.
Recommendation — Bind interactive agent actions to explicit authorisation checks before permitting execution.
NIST Zero Trust (SP 800-207)Principle of least privilege — Least privilegeInteractive agent sessions should only expose the minimum action scope required for the workflow.
Recommendation — Apply least privilege to agent sessions so embedded UI cannot expand effective access.
OWASP API Security Top 10API5 — Broken Function Level AuthorizationUI events can become unauthorised function calls if action-level checks are weak.
Recommendation — Enforce function-level authorisation on every action generated by embedded UI components.

Key terms

  • Interactive Agent Interface: A user-facing control surface that allows an AI agent to present, receive, and act on structured interactions rather than plain text alone. In governance terms, it expands the trust boundary because interface events can carry execution intent, not just display content.
  • UIResource: A UIResource is an MCP-UI object that packages interactive content for rendering inside an agent conversation. It carries a URI, MIME type, and content payload so the client can decide how to present the component while preserving protocol structure and security boundaries.
  • Structured UI Action: A typed event emitted by an embedded component that requests a tool call, expresses intent, sends a prompt, or triggers another controlled action. The important security property is that the event is machine-readable and policy-checkable before execution.
  • Sandboxed Rendering: A rendering pattern that isolates embedded content from the host application, usually through a restricted iframe or similar boundary. It reduces direct DOM and code exposure, but it does not by itself authorise what the embedded content may cause the agent to do.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity security are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 8, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org