By NHI Mgmt Group Editorial TeamBased on WorkOS: “Promptfoo vs. WorkOS: Security Testing Meets Enterprise Authentication” (November 3, 2025)

TL;DR: Promptfoo’s adversarial red-teaming probes target prompt injection, privilege escalation, memory poisoning, and goal hijacking in AI applications, while WorkOS provides the underlying authentication and authorization layer these agents rely on, according to WorkOS. The governance lesson is that validation and enforcement must be designed together, because testing alone cannot compensate for weak identity controls.


At a glance

What this is: This is an analysis of how AI agent security testing and enterprise authentication fit together, with the key finding that red-teaming can validate controls but cannot replace identity enforcement.

Why it matters: IAM, PAM, and NHI teams need to treat agent testing and runtime authorization as separate layers, because insecure identity design in AI systems cannot be fixed by test coverage alone.


Context

AI agent auth testing now sits between two different security problems: proving controls work and actually enforcing identity decisions at runtime. Promptfoo focuses on adversarial validation, while WorkOS is described as the authentication infrastructure that agents rely on for identity, authorization, and audit logging.

That distinction matters because AI agents combine application logic, tool use, and access decisions in one runtime path. When testing and enforcement are separated, teams can mistake a passing red-team result for governance maturity, even though the underlying authorization model may still be too broad or too brittle for production.

The primary identity issue here is not whether AI systems can be tested, but whether enterprise auth controls are strong enough to withstand prompt-driven abuse without depending on the test layer to do the job of runtime security.


Key questions

Q: What breaks when AI agents are given access without identity governance?

A: What breaks is accountability. The organisation may see actions, logs, and alerts, but it cannot reliably tie them to a governed identity with clear scope and revocation. That creates uncontrolled blast radius, especially when agents can reach sensitive systems through shared tokens, delegated service accounts, or broad API access.

Q: Why do AI agents create more authorization risk than static service accounts?

A: AI agents can vary their access needs by task, context, and timing inside the same workflow, which makes static entitlement assumptions weaker. If the control model assumes access is stable, it will either overgrant by default or block legitimate work. That is why fine-grained, real-time evaluation matters.

Q: How do security teams know if agent authorization is actually working?

A: Authorization is working only if the agent can complete the intended task without gaining unnecessary reach. Good signals include short-lived credentials, task-scoped permissions, approval for sensitive changes, and clear logs linking each action to a user and an agent. If credentials are reused, privileges persist, or the agent can move between systems without reauthorization, the control is failing.

Q: What is the difference between AI security testing and enterprise authentication?

A: AI security testing checks whether an attack can bypass controls. Enterprise authentication and authorization decide who can act, what they can reach, and how those decisions are enforced at runtime. Testing is validation, while authentication is the control plane that protects production access.


Technical breakdown

Why prompt-based attacks stress enterprise authorization

Prompt injection, privilege escalation, and goal hijacking are not just model-safety issues. In AI agents, they become identity problems when natural language input is used to influence tool calls, data access, or action selection. That means the security boundary is not the prompt alone, but the combination of user identity, application permissions, and downstream system authority. If those layers are loosely coupled, an attacker can exploit the agent’s decision path to reach resources the user was never meant to touch. The article’s key technical point is that adversarial testing can reveal these weaknesses, but only the authorization layer actually blocks them in production.

Practical implication: treat prompt injection as an authorization stress test, not as a substitute for access control design.

How trace-based testing complements runtime identity controls

Trace-based testing uses runtime instrumentation, such as OpenTelemetry, to inspect how an agent behaved during real or production-like execution. That is different from static review because it captures the full chain of calls across LLMs, retrieval systems, and tool interfaces. For identity teams, this matters because the most dangerous failures often appear only when a valid identity is combined with unexpected context, weak scope boundaries, or hidden privilege inheritance. The technical value is in correlating actual execution traces with the permissions that were in force at the moment of action.

Practical implication: use trace analysis to verify whether the identity layer and the agent’s observed behaviour align under adversarial conditions.

Why SSO, directory sync, and fine-grained authorization are separate controls

Enterprise authentication for AI agents is not a single control. SSO establishes who the user is, directory sync keeps role assignments current, and fine-grained authorization determines what the agent may do at the feature or API level. Those are different enforcement points, and collapsing them into one policy layer creates blind spots. In agentic systems, the dangerous pattern is letting a model decide policy in text while the real authorization check is too coarse, stale, or inconsistently applied. The article frames WorkOS as the runtime layer for those decisions, which is the right architectural distinction for identity governance.

Practical implication: separate authentication, entitlement synchronisation, and authorization enforcement in your agent architecture.


Threat narrative

Attacker objective: The objective is to make the AI agent act outside its intended authorization boundary and use that access to expose data or perform restricted actions.

  1. Entry occurs when an attacker reaches an AI agent through prompt injection or a crafted input path that the system accepts as legitimate.
  2. Credential or privilege abuse follows when the agent’s tool-calling logic or authorization checks can be steered beyond intended scope.
  3. Impact occurs when the agent leaks data, performs unauthorized actions, or exposes sensitive systems through the permissions it was allowed to use.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Validation and enforcement are separate security functions, not interchangeable ones. Promptfoo-style red-teaming can prove that a control fails under adversarial input, but it cannot create the control itself. That distinction matters for identity programmes because testing output is only evidence of resilience if the underlying auth model already exists. The practitioner conclusion is to design enforcement first and validation second.

Prompt injection becomes an identity issue when the agent can act on it. Once an AI system can call tools, read connected data, or trigger actions, the prompt is no longer just text. It is part of the authorization path, which means the security question shifts from content safety to permission scope, auditability, and boundary enforcement. Teams that still treat this as an application-only problem will understate the governance risk.

Agentic auth testing exposes an assumption gap between human-paced review and machine-paced action. Access review processes were designed for stable permissions that persist long enough to be certified. AI agents can request, use, and compound access within a session, so the review model sees the wrong state at the wrong time. The implication is that identity governance must move closer to issuance and runtime decisioning.

Runtime authorization is now a control plane for AI agents, not a back-office concern. SSO, directory sync, and fine-grained authorization determine whether the agent can safely execute in enterprise environments. In an agentic workflow, those controls define the boundary between a useful automation and an uncontrolled actor. Practitioners should treat authorization architecture as part of the product design, not a later compliance step.

Identity blast radius is the right named concept for this category. The more an AI agent can chain identity, tool use, and data access into one session, the larger the blast radius of a single authorization weakness. That is why testing, logging, and policy enforcement have to be connected. The practitioner takeaway is to measure how far one authenticated agent action can propagate before a human or policy boundary stops it.

From our research library:

What this signals

Identity blast radius: AI agent governance should be judged by how far one authenticated action can propagate across tools, data, and workflows before it hits a hard boundary. That makes authorization architecture a programme-level concern, not just an application integration task.

Testing can still miss the most damaging failure mode: a control that looks correct in a red-team report but is too loosely bound to the runtime identity layer to stop real abuse. Teams should focus on whether permissions are enforced where the action occurs, not where the test is written.


For practitioners

  • Define the agent’s authorization boundary before testing Map every user, role, tool, and data source the AI agent can touch, then assign explicit limits to each permission path before running red-team probes.
  • Separate authentication from agent decision logic Keep SSO, directory sync, and fine-grained authorization in the runtime identity layer so prompt content never becomes the source of truth for access decisions.
  • Instrument agent actions with trace-level logging Capture tool calls, policy decisions, and data access in a form that lets you reconstruct whether the agent stayed inside its approved scope during an attack.
  • Red-team the controls, not just the model Use adversarial testing to validate that authorization, isolation, and logging survive prompt injection, role manipulation, and privilege escalation attempts.

Key takeaways

  • AI agent security testing is useful, but it does not replace the runtime identity controls that actually enforce who can do what.
  • The strongest risk signal in this article is the mismatch between adversarial validation and enterprise authorization enforcement.
  • Practitioners should map agent permissions, instrument runtime decisions, and test the access boundary as a single control surface.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseThe article centers on AI agents being probed for privilege escalation and authorization bypass.
ASI09 — Human-Agent Trust ExploitationPrompt-driven manipulation can cause users or systems to trust an agent beyond its safe authority.
Recommendation — Map agent privilege paths to ASI03 and validate that tool access cannot exceed assigned scope. Assess where users can be misled into granting an agent actions outside its intended trust boundary.
NIST AI RMFGOVERN — AI Governance and AccountabilityThe article is about how AI agent controls and testing fit into governance, accountability, and oversight.
Recommendation — Use GOVERN to assign ownership for agent identity controls and validation outcomes.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe core issue is whether AI agents’ access rights are enforced correctly at runtime.
Recommendation — Apply PR.AA-05 to keep agent entitlements explicit, current, and enforced at the point of access.
OWASP Non-Human Identity Top 10NHI-05 — Overprivileged NHIAI agents operate as non-human identities when they are granted access to enterprise systems.
Recommendation — Review AI agent permissions for overprivilege and shrink access to the minimum tool scope.

Key terms

  • Agentic Access: Agentic access is delegated system access granted to an AI agent or autonomous workflow so it can perform defined tasks across tools and data sources. It differs from human access because the actor can execute continuously, combine actions quickly, and amplify mistakes at scale.
  • Authorization Boundary: The authorization boundary is the defined scope of systems, identities, and dependencies that must satisfy a compliance programme. In FedRAMP, it determines what the assessor evaluates and what must be documented as external, so boundary accuracy is a control decision, not a paperwork exercise.
  • Trace-Based Evaluation: An evaluation approach that records the full execution path of a run, including inputs, intermediate calls, retrieved context, and outputs. It helps teams debug multi-step AI systems by showing how a result was produced, not just whether it looked correct.
  • Identity Blast Radius: The amount of damage a compromised identity can cause across systems, data, and infrastructure. In NHI environments, it is shaped by permissions, network reach, and administrative capability rather than by the credential alone. Reducing blast radius is a containment strategy that limits lateral movement and data exposure.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org