By NHI Mgmt Group Editorial TeamBased on WorkOS: “Anthropic’s Computer Use versus OpenAI’s Computer Using Agent (CUA)” (July 30, 2025)

TL;DR: Anthropic’s Computer Use and OpenAI’s Computer Using Agent point to a new class of AI that can act inside desktop and browser environments, exposing identity controls that were built for predefined access rather than runtime action selection, according to WorkOS. Access review models assume a stable identity with predictable scope, but these agents can choose actions across arbitrary software within a task.


At a glance

What this is: This analysis compares Anthropic’s Computer Use and OpenAI’s Computer Using Agent and concludes that computer-use agents create an identity control gap because they act across desktop and browser software in ways traditional access models do not anticipate.

Why it matters: IAM, NHI, and PAM teams need to account for agent runtime behaviour, because task-scoped UI action is not the same as static entitlement and can widen the governance gap between access granted and access exercised.


Context

Computer use agents are AI systems that interact with desktops or browsers through a user interface instead of only through APIs. That matters for identity governance because the control question changes from who can call a service to what an agent is allowed to do across arbitrary software during a task.

WorkOS frames Anthropic’s Computer Use and OpenAI’s Computer Using Agent as early examples of this shift. The article’s core concern is not model quality alone, but whether current IAM and NHI assumptions can still govern software that selects actions at runtime inside managed or direct-access environments.


Key questions

Q: What breaks when computer-use agents are given broad desktop access?

A: Broad desktop access breaks the assumption that identity scope can be enumerated in advance. The agent can move across native apps, files, and web interfaces during a single task, so the real control boundary is the execution environment, not the login event. Without that boundary, privilege review and containment both become weaker.

Q: Why do computer-use agents create more risk than ordinary workflow automation?

A: They create more risk because the sequence of actions is chosen at runtime, not pre-scripted as a fixed workflow. That makes it harder to predict which applications, screens, or data paths the agent will touch, and it makes policy based only on credentials or API scope incomplete.

Q: What signs show that a computer-use agent is operating outside its intended boundary?

A: Look for unexpected application switching, repeated retries across unrelated screens, and task completion that depends on software the original request never mentioned. Those signals suggest the agent is generalising beyond the intended control envelope and that the current access model is too coarse.

Q: Should organisations use browser-only agents or direct desktop agents for sensitive work?

A: Browser-only agents are easier to contain and observe, so they are usually the safer choice when the use case can be satisfied in the web layer. Direct desktop agents should be reserved for workflows that truly require native applications, local files, or system-level interaction.


Technical breakdown

Desktop control versus browser-bounded execution

Anthropic’s Computer Use gives a model direct control over the desktop, so the agent can interact with native applications, system tools, and websites through screenshots, clicks, and keystrokes. OpenAI’s Computer Using Agent keeps the model inside a secure virtual browser, which narrows the operating surface to web-based tasks. That architectural split is not cosmetic. Desktop reach increases functional flexibility, while browser confinement reduces the range of software the agent can touch but also leaves it dependent on web UI structure. In identity terms, the control boundary is no longer just authentication to a service. It becomes the runtime context in which an agent is permitted to operate.

Practical implication: Practitioners should classify whether an agent needs desktop-level access or only browser-bounded execution before assigning any identity scope.

Why UI agents strain least privilege and access review

Traditional least privilege assumes the actor’s permitted actions can be bounded in advance. UI-driven agents break that assumption because the task is expressed at a high level and the exact sequence of clicks, fields, and applications emerges only at runtime. That means privilege is exercised through interface paths that are hard to enumerate upfront, and the same permission can reach multiple systems in one workflow. Access review becomes weaker when the governance record says one thing but the agent can traverse many interfaces to get work done. This is why computer-use agents create an identity control gap rather than just a new automation channel.

Practical implication: Security teams should treat task scope, UI reach, and allowed applications as part of the access decision, not as implementation detail.

Sandboxed virtual environments versus direct machine access

A managed virtual browser or sandbox changes the threat and governance model because it contains the agent’s environment, network reach, and software surface. Direct machine access, by contrast, places the agent closer to local applications, files, and user context, which raises the chance of unintended cross-application effects. The difference matters for identity because confinement is one of the few practical ways to make an agent’s runtime behaviour auditable and reversible. A sandbox does not solve authorisation by itself, but it creates a narrower blast radius and a clearer boundary for monitoring, logging, and termination. Direct desktop control offers more capability, but it also creates more governance ambiguity.

Practical implication: Choose the most constrained execution environment that still meets the use case, then map monitoring and termination controls to that boundary.


Threat narrative

Attacker objective: The objective is not necessarily theft in the article’s framing, but autonomous completion of multi-step work across environments that traditional identity controls were not designed to constrain.

  1. Entry occurs when a computer-use agent is granted permission to operate inside a desktop or browser environment on behalf of a user or workflow.
  2. Escalation happens when the agent uses that permission to move beyond a single action and execute a multi-step sequence across native apps, files, or web interfaces.
  3. Impact follows when the agent completes sensitive work, modifies records, or touches systems that were not originally scoped as explicit API targets.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Computer-use agents create an identity control gap because task intent is not the same as execution scope. The article shows that a high-level instruction can unfold into clicks, navigation, typing, and cross-application state changes that no provisioning-time policy fully describes. That is a governance problem, not just a product-design problem. Practitioners should treat the gap between intent and runtime behaviour as the central control question.

Least privilege becomes harder to define when the actor decides the path at runtime. The article’s comparison between direct desktop control and browser-bounded execution shows that the same task can produce very different access paths. Existing IAM and NHI models assume the path can be scoped in advance, but computer-use agents discover the path while working. The implication is that entitlement design must account for agent-timed execution, not just named permissions.

Managed confinement is the closest thing to a governance boundary for computer-use agents. The article’s virtual browser model is not a complete answer, but it is a stronger control shape than unrestricted desktop reach because it narrows the software surface and improves auditability. That makes environment design part of identity governance, not just infrastructure hygiene. Practitioners should think in terms of execution boundary first, control catalog second.

Runtime autonomy changes how access review, accountability, and incident response fit together. When an agent can complete a task by selecting actions during the session, the traditional assumption that access will persist long enough to be reviewed starts to erode. Who approved the task, who owned the workflow, and what the agent actually did become intertwined. Identity governance therefore has to move closer to runtime authorisation and observability, not rely on after-the-fact certification alone.

Computer-use agents should be treated as governed non-human identities, not as generic automation. The article is about AI systems that can operate software on behalf of people, which places them in the NHI category with added autonomy pressure. That distinction matters because the control set expands from secrets and service accounts into execution policy, environment confinement, and auditability. Practitioners should govern the agent, not just the credentials it uses.

What this signals

Task scope is becoming an identity control problem. Computer-use agents make the gap between what a workflow is supposed to do and what the agent can actually do much more visible. For practitioners, that means entitlement design has to incorporate execution boundary, application reach, and session observability together, not as separate programmes.

Runtime behaviour matters more than model label. The important distinction is not whether the system is called an agent, a copilot, or an assistant, but whether it can select actions during execution without predetermined steps. Identity programmes that still classify access only by credential type will miss the governance shift created by autonomous UI action.

Confinement is now part of the identity discussion. Managed virtual environments reduce ambiguity because they narrow the agent’s operating surface and give security teams a clearer place to log, monitor, and terminate activity. That makes environment design a first-class control decision for NHI and agentic AI governance.


For practitioners

  • Define agent execution boundaries Specify whether the agent may use a browser, a desktop, or both, and map that choice to the smallest feasible software surface and network reach.
  • Separate task approval from access scope Record who approved the task, which applications are in scope, and which actions remain prohibited even when the agent is operating legitimately.
  • Instrument session-level observability Log screenshots, action sequences, and application transitions so you can reconstruct what the agent did inside a session and where it crossed boundaries.
  • Restrict direct machine access where possible Prefer managed virtual environments for high-risk workflows and reserve direct desktop control for cases that genuinely require native application access.
  • Review NHI governance for agentic workflows Extend identity lifecycle, authorisation, and offboarding checks to the agents themselves, not only to the human operators and service accounts around them.

Key takeaways

  • Computer-use agents blur the line between access and execution because they can choose actions across applications during a task rather than follow a fixed integration path.
  • The governance gap is not just technical; it is structural, because existing IAM assumptions about predefined scope and reviewable privilege no longer fit runtime UI control.
  • The strongest near-term control is to constrain the execution boundary, then add observability and task-scoped authorisation around that boundary.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseComputer-use agents turn runtime action scope into the central control issue.
Recommendation — Constrain agent identity and privilege to the smallest execution boundary that still supports the task.
OWASP Non-Human Identity Top 10NHI-05 — Overprivileged NHIDesktop-capable agents can exercise far more access than a static entitlement suggests.
NHI-08 — Environment IsolationThe article contrasts managed virtual environments with direct machine access.
Recommendation — Map agent workflows to least-privilege scopes and remove access paths the task does not require. Isolate high-risk agent sessions in managed environments that narrow software and network reach.
MITRE ATT&CKTA0008 — Lateral MovementA UI agent can traverse multiple applications during a single session.
Recommendation — Trace agent activity across applications to detect unexpected cross-system movement during one workflow.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe core issue is whether permissions match what the agent can do at runtime.
Recommendation — Align entitlements with approved task boundaries and verify them before agent execution starts.

Key terms

  • Computer-use agent: An AI system that can observe a user interface and take actions across software on behalf of a task. In practice, it extends identity governance beyond API access because the agent can navigate live applications, combine steps, and adapt to changing state during the session.
  • Execution boundary: The point at which an authorised task turns into a real system change, such as writing data, deleting records, spending money, or invoking a downstream tool. In AI governance, controlling the execution boundary matters more than simply approving access, because harm occurs when actions are allowed to complete unchecked.
  • Managed virtual environment: A contained runtime such as a sandboxed browser or desktop where the agent operates under tighter network and software limits. This reduces blast radius and improves auditability, but it still requires explicit governance over task scope, logging, and termination conditions.
  • Runtime Authorisation: Runtime authorisation is the practice of deciding access while a task is in progress, rather than only at provisioning time. It matters for NHIs because credentials and entitlements can change risk mid-session, especially when automation or AI agents interact with sensitive systems.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 8, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org