By NHI Mgmt Group Editorial TeamBased on WorkOS: “Enterprise AI Agent Playbook: What Anthropic and OpenAI Reveal About Building Production-Ready Systems” (July 23, 2025)

TL;DR: Anthropic and OpenAI’s enterprise guidance shows that production AI agents succeed with simple composable patterns, layered guardrails, and explicit tool-risk controls, while enterprise teams still struggle with evaluation, security, and delegation across systems, according to WorkOS’s analysis of the two guides. The deeper issue is that traditional IAM assumes stable, reviewable access, but agentic systems can act, branch, and delegate within one session.


At a glance

What this is: This analysis contrasts enterprise AI agent playbooks with the operational and security gap still facing most teams: production agents need simple patterns, layered guardrails, and explicit tool-risk controls, not just more model capability.

Why it matters: IAM and security leaders need to treat AI agents as systems that create, route, and exercise access across tools and sessions, which changes how delegation, authorisation, and auditability have to be governed.


Context

Enterprise AI agents are not just another automation layer. They can choose actions, call tools, and complete workflows on a user’s behalf, which means the governance problem is no longer only model quality but delegated execution across systems.

This article sits squarely in the AI agent identity and access problem space. The production gap is not whether agents can reason, but whether enterprises can constrain tool use, evaluate behaviour, and keep authorisation auditable when the workflow changes at runtime.

WorkOS uses Anthropic and OpenAI playbooks to show that successful deployments rely on composable patterns, layered guardrails, and tool-risk awareness, while many teams still treat agent build-out like ordinary application development.


Key questions

Q: What should organisations do before deploying AI agents in enterprise workflows?

A: Define the agent’s identity, privilege scope, and accountability before enabling production access. Then add output validation for harmful or non-compliant responses. That sequence gives security, IAM, and compliance teams a clear chain of evidence when the agent touches regulated or customer-facing data.

Q: Why do AI agents create more risk than traditional automation?

A: AI agents create more risk because they can interpret context, choose actions, and invoke tools autonomously. Traditional automation follows fixed rules, but an agent can be manipulated into using its own authority in unintended ways. That makes permission scope, tool boundaries, and monitoring more important than model accuracy alone.

Q: What are the signs that an AI security model is failing or becoming unreliable?

A: Common warning signs include rising false positives, missed threats, inconsistent outputs, and recommendations that security teams cannot explain or validate. Poor results often point to weak training data, stale models, or poisoned inputs. When analysts spend more time correcting the model than using it, the system is no longer improving security and may be creating operational noise.

Q: How should security teams govern agent access to headless enterprise systems?

A: Security teams should govern agent access by treating APIs, tools, and protocols as runtime identity surfaces. That means binding authorization, audit, and rate limits to each request, not just to the application. Teams should also scope context tightly, because an agent that can retrieve too much data can do damage even when its credentials are valid.


Technical breakdown

Why simple composable agent patterns outperform complex stacks

Both playbooks converge on a practical architectural point: most production value comes from simple, composable workflows rather than sprawling multi-agent systems. Prompt chaining, routing, parallelisation, orchestrator-workers, and evaluator-optimizer patterns work because they isolate uncertainty, bound the decision surface, and make failures observable. The key design idea is not more autonomy, but better decomposition. That matters because enterprise AI systems often fail when a single prompt tries to do too much, or when orchestration logic is hidden inside an opaque stack. The more explicit the workflow, the easier it is to test, reason about, and govern.

Practical implication: prefer bounded workflow patterns that can be evaluated step by step before granting broader tool access.

Layered guardrails are stronger than single-point content filtering

The article describes a defence-in-depth model where LLM-based guardrails, rules-based filters, moderation layers, relevance classifiers, and safety classifiers each catch different failure modes. That structure matters because prompt injection, indirect instruction attacks, and harmful content are not the same problem. A regex cannot understand intent, and a model guardrail cannot reliably catch every known pattern. In production, guardrails need to operate across the input, output, and session context, not just at a single request boundary. The operational lesson is that agent safety is a system property, not a single control.

Practical implication: build layered controls that combine deterministic checks with context-aware classification and session-level review.

Tool risk changes the identity problem for enterprise agents

WorkOS separates tools into data, action, and orchestration categories, which is the right way to think about identity exposure in agentic systems. Read-only tools expose information risk, action tools create durable business change, and orchestration tools can propagate permission through chained agent behaviour. That means authorisation cannot stop at model access. It has to define which tool classes an agent may invoke, under what conditions, and with what audit trail. Once an agent can write to systems, the security model shifts from information retrieval to delegated execution control.

Practical implication: classify every agent tool by risk tier before enabling access, especially where writes or downstream delegation are possible.


Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Production AI agents expose an access model, not just an automation model. The article shows that the real break from traditional software is not speed or scale, but delegated action across systems with different trust boundaries. Once an agent can choose tools and execute work on behalf of a user, identity governance has to account for runtime decision paths, not only provisioned entitlements. The practitioner conclusion is that agent access is an execution model that must be governed as such.

Simple workflow patterns are a governance advantage, not just an engineering preference. Anthropic and OpenAI both converge on composable patterns because they make behaviour easier to evaluate, isolate, and recover. That is important for governance teams because a bounded workflow creates fewer hidden states than a large multi-agent design. The practitioner conclusion is that complexity should be justified by measurable improvement, not architectural novelty.

Tool-risk classification should become part of authorisation design. The article’s data, action, and orchestration split is more useful than a generic allow-list because it mirrors business impact. Read-only access, write access, and agent-to-agent chaining do not belong in the same control bucket. The practitioner conclusion is that authorisation policy must follow tool risk, not just application role.

Agent security is now an identity lifecycle problem as much as an application security problem. The production gap sits in evaluation, delegation, and auditability because agents can take action over time and across systems. That means access reviews, session controls, and offboarding logic need to reflect how an agent operates rather than how a human user logs in. The practitioner conclusion is that lifecycle governance has to extend into agent behaviour, not stop at account creation.

Named concept: tool-risk posture. Production readiness depends on whether teams can distinguish informational tools from state-changing tools and set controls accordingly. That distinction is what separates low-consequence retrieval from delegated business action. The practitioner conclusion is to govern agents by tool-risk posture before broadening their operating scope.

From our research library:

What this signals

Tool-risk posture: The next governance step for enterprise teams is to stop thinking about agents as a single identity class and start classifying their tool access by risk. Read-only retrieval, state-changing actions, and downstream orchestration should not share the same approval model or review cadence.

Production programmes will struggle if they keep evaluation separate from authorisation. Agent behaviour has to be tested under realistic conditions before access expands, because the failure mode is not just bad output, but uncontrolled action across systems.

Access reviews assume a stable entitlement window, but agentic workflows can create and consume authority inside one business process. That makes issuance-time control more important than retrospective certification for high-risk tools.


For practitioners

  • Define agent use cases by decision complexity Start with workflows that involve nuanced judgment, exception handling, or unstructured data rather than tasks already suited to traditional automation. If a deterministic process can do the job, keep the agent out of the path.
  • Classify tools by risk tier Separate read-only data tools from write-capable action tools and downstream orchestration tools, then set authorisation boundaries and audit expectations for each class.
  • Add layered guardrails before production rollout Use deterministic rules, context-aware classifiers, and session-level checks together so that prompt injection, unsafe content, and scope drift are not handled by one control alone.
  • Instrument evaluation as part of deployment Measure whether prompts, routing, and tool use improve outcomes under realistic conditions, and treat evaluator feedback as a mandatory deployment signal rather than a nice-to-have.
  • Govern agent delegation like privileged access Track which systems an agent can reach, what it can change, and how long that authority persists, then align review and offboarding with the actual operating model.

Key takeaways

  • AI agents change the governance problem by combining decision-making, tool use, and delegated action in one runtime workflow.
  • The production gap is less about model sophistication than about evaluation discipline, layered guardrails, and explicit tool-risk boundaries.
  • Enterprises need authorisation and lifecycle controls that match how agents actually operate across systems, not how human users log in.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI02 — Tool MisuseThe article centres on how agents choose and invoke tools in production workflows.
ASI03 — Identity & Privilege AbuseDelegated access and runtime action make privilege control central to the article.
Recommendation — Restrict agent tool use by risk tier and monitor for misuse across tool boundaries. Bind agent privileges to explicit delegation rules and audit every privilege-bearing action.
NIST AI RMFMANAGE — AI Risk ManagementThe article is fundamentally about managing operational AI risk in production.
Recommendation — Embed AI risk controls into deployment, monitoring, and change management for agent systems.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsAgent access to enterprise systems depends on tightly governed permissions and entitlements.
Recommendation — Apply PR.AA-05 to define, limit, and review agent permissions by system and action type.
OWASP Non-Human Identity Top 10NHI-05 — Overprivileged NHIAI agents act as non-human identities and the article warns against broad, uncapped access.
Recommendation — Reduce agent blast radius by removing excess permissions from every non-human identity.

Key terms

  • Agentic AI: Autonomous AI systems capable of planning, deciding, and taking actions, including calling APIs, writing code, and orchestrating other agents, with minimal human oversight. Agentic AI introduces new NHI risks as agents must authenticate to external services.
  • Tool Risk: Tool risk is the security impact created by what an AI agent is allowed to do through external functions or APIs. Read-only tools, write-capable tools, and orchestration tools carry different blast radii, so access policy has to distinguish them rather than treating all tools equally.
  • Delegated Execution: Delegated execution is when software is allowed to perform actions on behalf of a user, process, or business function. In NHI governance, the risk is that the delegated actor may chain actions beyond the original intent, so controls must focus on scope, approval, and revocation.
  • Guardrail Layering: Guardrail layering is the practice of combining multiple independent controls so that one failure does not expose the full system. In AI security, that usually means pairing cloud configuration controls, model behaviour checks and identity restrictions across the same workflow.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 8, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org