Join our Newsletter — 33% off our NHI Course

How should teams implement code mode for AI agents when the agent must operate inside a live web application that a user is also using?

Teams should expose only the browser actions the user can already perform, wrap them in a typed SDK, and execute agent code in a Web Worker rather than the main thread. That keeps the agent aligned with UI state, preserves responsiveness, and limits blast radius. The key control is the browser action catalog, because it turns interface behavior into governed capabilities the agent can call safely.

How code mode should be constrained inside a live user session

Code mode works best when the agent is treated as a bounded UI participant, not a second hidden operator. The agent should only be able to trigger the same browser actions a user could perform, and those actions should be exposed through a typed SDK so every capability is explicit, reviewable, and easy to validate against the live interface state.

That design matters because a live web app is stateful: the user may be editing, navigating, or confirming something at the same time the agent is acting. If the agent can reach into the DOM or call arbitrary page functions, it can drift out of sync with what the user sees and create confusing or destructive side effects.

The practical rule is to define the action catalog first, then map agent intent to those actions, rather than letting the model improvise against the page. If the browser action is not something a human could reasonably do in that moment, it should not be exposed as an agent capability.

Why the browser action catalog is the control that keeps the agent governable

The browser action catalog is the main control because it turns interface behavior into a capability boundary. A cataloged action can be named, typed, logged, and versioned, which makes it much easier to reason about what the agent is allowed to do than a free-form page scripting approach. That is the difference between governed automation and opaque UI manipulation.

This also creates a cleaner authorization model for the agent. Rather than granting broad page access, teams can decide which actions are available in which states, and whether an action needs extra confirmation when it becomes high impact. The same pattern is useful in live workflows where the user and agent share the same session but should not share unlimited control.

Good catalog design also forces teams to think about state transitions, not just buttons. An action such as “submit form” may be safe only after validation, draft saving, or user review has happened. Encoding those preconditions in the SDK keeps the agent aligned with the real workflow instead of the page markup.

Why execution belongs in a Web Worker, not the main thread

Running the agent code in a Web Worker protects the user experience and reduces the blast radius of mistakes. The main thread should remain dedicated to rendering, input handling, and direct user interaction. If the agent blocks that thread, a slow model call, a loop, or a bad tool invocation can freeze the app at the exact moment the user needs control.

A worker also gives teams a more defensible isolation boundary for orchestration logic. The agent can compute plans, process results, and queue approved browser actions without sharing the same execution path as the UI event loop. That makes cancellation, throttling, and interruption easier when the user takes over or the page state changes underneath the agent.

Teams should still treat the worker as privileged application code, not as a sandbox that removes all risk. The worker is a containment measure for responsiveness and fault isolation, while the browser action catalog is the actual governance mechanism that limits what the agent may do.

Risk and Threat Considerations

Live-session code mode creates exposure when the agent can act faster than the user can notice. The main failure pattern is overreach, where a model-generated plan executes against stale UI state or uses a capability that is broader than the visible interface suggests. That can lead to unintended writes, destructive actions, or user confusion during concurrent activity.

Failure mechanism: If the agent is allowed to call arbitrary page functions, infer hidden state, or bypass the action catalog, it can operate outside the user’s mental model and outside the browser’s visible controls.

Impact: Teams can see incorrect submissions, broken workflow sequencing, lost user edits, and hard-to-audit actions that are difficult to roll back once the session state diverges.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5, OWASP ASVS and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
OWASP Agentic AI Top 10 ASI03 — Identity & Privilege Abuse Code mode inside a live user session must constrain agent authority and action scope.
Recommendation — Restrict agent actions to least-privilege capabilities and require approval for high-impact steps.
NIST SP 800-53 Rev 5 AC-6 — Least Privilege The browser action catalog limits the agent to only permitted interface capabilities.
SC-39 — Process Isolation Running agent code in a Web Worker is an isolation choice that reduces UI blast radius.
Recommendation — Limit agent-exposed browser actions to the minimum needed for the task. Isolate agent execution from the main UI thread to reduce impact from faults or abuse.
OWASP ASVS V8 — Authorization Typed, governed browser actions are an authorization boundary for agent-driven UI operations.
Recommendation — Authorize each agent action explicitly against the current UI state before execution.
NIST CSF 2.0 PR.AA-05 — Least Privilege The answer centers on minimizing agent power and constraining what it can do in-session.
Recommendation — Enforce least-privilege access for agent-controlled browser capabilities.

Practitioner Guidance

What to verify: Ensure every agent-exposed action has a human-readable name, explicit input types, and a clear state precondition. If a browser action cannot be explained as a normal user action in the current screen state, do not expose it.

Decision rule: If the action can change durable data, advance a workflow, or trigger external effects, require an explicit approval point or a stricter policy gate before the worker can invoke it. Reserve silent auto-execution for low-impact UI operations.

What good looks like: The user can pause, correct, or override the agent without losing control of the page, and the team can reconstruct exactly which cataloged actions were called and why.

Practitioner takeaway: The safest code mode is not the most capable one, it is the one where the agent’s reach matches the user’s visible authority and the browser remains the source of truth for what can happen next.