By NHI Mgmt Group Editorial TeamBased on WorkOS: “Haize Labs: AI Safety Testing” (November 7, 2025)

TL;DR: Haize Labs’ red-teaming platform is built to find prompt injection, goal misalignment, hallucination and other behavioral failures in LLMs and AI agents, while WorkOS positions enterprise authentication as a separate control plane for access and user management. Behaviour testing can prove an AI system is unsafe, but it cannot substitute for identity, authorization or lifecycle governance.


At a glance

What this is: This is an analysis of AI safety testing versus enterprise authentication, arguing that red-teaming AI behaviour and governing access are separate controls with different failure modes.

Why it matters: IAM, IGA and PAM teams need to separate AI behavioural validation from access governance so they do not mistake model safety tooling for identity control.

By the numbers:

  • WorkOS says Haize Labs received a $100M post-money valuation.
  • Cascade delivered 38x faster attack generation with 4x reduction in GPU memory usage.

Context

AI agent safety testing focuses on behavioural failure, not access control. In this article, the core issue is whether an AI system can be pushed into unsafe output through prompt injection, goal misalignment or adversarial inputs, and why that problem sits alongside, not inside, enterprise authentication.

For IAM and identity governance teams, the distinction matters because an AI application can be authenticated correctly and still behave unsafely after login. The inverse is also true: safety testing may expose model weaknesses, but it does not grant, revoke or scope access for users, admins, service accounts or AI agents.

WorkOS uses Haize Labs as the example to separate two control planes that enterprises often blur together: behavioural assurance for the model, and identity assurance for the people and systems around it. That split is now central to AI governance programmes that have to support production use, auditability and least privilege.


Key questions

Q: How should security teams test AI systems for safety and security separately?

A: Run two evaluation tracks. Safety tests should measure harmful content, bias, refusal quality, and policy compliance. Security tests should focus on prompt injection, hidden instructions, tool misuse, data leakage, and unauthorised actions. If a system is only tested for one dimension, teams can ship a model that sounds safe while still being easy to manipulate.

Q: Why can enterprise authentication still leave AI agents unsafe?

A: Because authentication proves identity, not behaviour. An AI agent can be correctly logged in through SSO or directory sync and still follow a malicious prompt, hallucinate sensitive content or ignore policy once inside the session. Access control reduces unauthorized entry, but it does not constrain how the model reasons or responds after access is granted.

Q: What are the signs that AI safety testing is being used as a proxy for access control?

A: The clearest sign is when teams point to red-teaming results, content filters or model evaluation dashboards as evidence that user access, entitlement scope or tenant isolation is acceptable. That is a category error. Safety testing exposes behavioural risk, while access control exposes identity and privilege risk. If the artefacts are being used interchangeably, governance is blurred.

Q: What should organisations do when AI access and AI safety are owned by different teams?

A: They should define a shared governance model with separate control objectives. The identity team should own provisioning, authorization, auditability and revocation, while the AI assurance team should own adversarial testing, runtime monitoring and failure-mode analysis. Common reporting helps, but the controls should not collapse into one programme because they answer different risk questions.


Technical breakdown

Behavioral red-teaming versus access control

Behavioral red-teaming is adversarial testing that tries to make an AI system fail through inputs, prompts or conversation flow. It is aimed at model behaviour, not identities or entitlements. Enterprise authentication, by contrast, decides who can reach the application, what they can see, and how access is provisioned or revoked. Those controls operate before and around the model, while red-teaming operates against the model itself. The two layers are complementary, but they answer different questions and produce different evidence. One finds unsafe responses or jailbreak conditions. The other establishes whether access is correctly authorised and governed across users, admins and connected systems.

Practical implication: Treat AI safety findings as model-risk evidence, not as proof that access controls are adequate.

Why continuous AI testing scales differently from IAM review

The article highlights automated red-teaming because manual testing does not scale to frequent model updates, prompt changes or retrieval changes. That makes sense for behavioural assurance, where each new version can introduce new failure modes. IAM review is different: it governs identity state, entitlement scope and lifecycle events, which are assessed through provisioning, access review and offboarding processes. Continuous testing can tell you whether a model now hallucinates or follows a malicious prompt. It cannot tell you whether the right user has the right role, whether directory sync is accurate, or whether a privileged account was removed on time. The evidence and operating rhythm are different by design.

Practical implication: Map AI safety testing into model release gates, and keep identity certification, provisioning and revocation in the IAM programme.

Enterprise auth is not an AI safety control

The article draws a hard line between enterprise authentication features and AI safety features. SSO, SCIM, admin portals and RBAC govern access to the application and its tenants; they do not evaluate whether the agent will follow a malicious instruction, leak context, or produce harmful output. That distinction matters because teams sometimes assume that once an AI product supports SSO or directory sync, the deeper safety problem is also covered. It is not. Authentication tells you who the requester is. Safety testing tells you how the system behaves once that requester is inside. Those are separate assurance questions with separate owners and evidence streams.

Practical implication: Do not use federation or provisioning features as a proxy for agent safety assurance.


Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

AI behavioural safety and enterprise identity are parallel control problems: this article is useful because it separates model red-teaming from access governance. Behavioural testing can show that an AI system is vulnerable to prompt injection or misalignment, but that evidence does not govern who can enter the system or what entitlements they receive. Practitioners need to stop treating model assurance as a substitute for IAM, because the failure modes, evidence and owners are different. The implication is that AI programmes need two assurance tracks, not one blended control narrative.

Enterprise authentication solves the wrong problem if it is treated as AI safety: SSO, directory sync and RBAC establish tenant access, but they do not bound the behaviour of an authenticated agent. That distinction becomes critical when enterprises deploy AI into production workflows, because a logged-in agent can still hallucinate, obey malicious prompts or mishandle sensitive context. The practical conclusion is that identity controls reduce unauthorized access, while red-teaming reduces behavioural uncertainty; neither collapses into the other.

Behavioural testing should be treated as model-risk evidence, not IAM evidence: the article shows why automated red-teaming belongs in deployment governance, not in the access review process. Access reviews assume a stable entitlement set that can be certified or revoked. Safety testing produces a different artefact: proof that a model failed under certain conditions. Those outputs may inform go-live decisions, but they do not validate identity lifecycle, privileged access or tenant isolation.

AI governance will split into assurance layers as production adoption grows: enterprises are now forced to govern AI systems the way they govern complex platform estates, with separate controls for identity, runtime behaviour and lifecycle accountability. That is a healthier model than asking one product class to solve both access and safety. The organisations that separate these layers will move faster because they can assign each risk to the right control owner and evidence stream.

Prompt injection is a model failure mode, not an access model replacement: the most important lesson here is that adversarial inputs exploit the system after access has already been granted. That means the security question changes from who can log in to what the system does with input once logged in. Practitioners should treat prompt injection testing as a runtime assurance input, not as a reason to relax IAM or authorization design.

From our research library:

What this signals

Identity review cadences do not solve behavioural risk: the article’s core distinction is that access controls answer who may use an AI system, while red-teaming answers what the system does once used. Organisations that blur those layers will overstate the value of authentication and understate the need for adversarial testing.

AI safety evidence belongs in model governance, not access certification: access reviews assume entitlements persist long enough to be reviewed, but the article’s testing model is about runtime behaviour under attack. The operational boundary should be clearer in production programmes that now combine AI features with enterprise login flows.

Access scoping remains the decisive identity variable for AI systems: according to the 2026 Infrastructure Identity Survey, systems with least-privileged AI access had a 17% incident rate vs 76% for over-privileged systems, and organisations failing to scope AI access properly are 4.5x more likely to experience a security incident. That is a governance signal to separate behavioural testing from privilege design, not to conflate them.


For practitioners

  • Separate model-risk testing from access governance Assign AI red-teaming, prompt injection testing and hallucination evaluation to the AI assurance process, and keep identity proofing, entitlement scoping and revocation under IAM ownership.
  • Gate production AI releases on behavioural evidence Require automated red-teaming results before model updates, prompt changes or retrieval changes are promoted into production, so safety testing becomes part of release qualification.
  • Keep enterprise auth evidence distinct Document SSO, SCIM, RBAC and audit logs as access controls, not as evidence that an AI agent behaves safely under adversarial input.
  • Review runtime prompts and retrieval paths Test the combinations of prompts, context windows and retrieval sources that can change model behaviour after authentication, because post-login abuse is a separate risk from access abuse.
  • Map control ownership across AI and identity teams Make sure the team that owns AI safety metrics is not also assumed to own tenant access, privileged administration or lifecycle governance for users and connected systems.

Key takeaways

  • This article shows that AI red-teaming and enterprise authentication solve different problems, and only one of them governs access.
  • WorkOS cites a $100M post-money valuation for Haize Labs and a 38x speedup in attack generation, which shows the market is investing in scalable behavioural testing.
  • The operational takeaway is to keep AI safety evidence, identity governance and privileged access controls in separate control paths.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI03 — Identity & Privilege AbuseThe article contrasts AI behaviour testing with the access and privilege boundaries around agents.
ASI09 — Human-Agent Trust ExploitationPrompt injection and unsafe outputs exploit user trust in agent responses, not just login paths.
Recommendation — Model agent privilege separately from model behaviour and test both before production release. Validate how prompts and user trust can be abused before allowing agents into production workflows.
OWASP Non-Human Identity Top 10NHI-04 — Insecure AuthenticationThe article explicitly separates enterprise authentication from AI safety, making auth correctness central to deployment.
NHI-05 — Overprivileged NHIThe article’s access discussion centres on scoping AI system access and avoiding excessive entitlements.
Recommendation — Use strong authentication for AI applications, but do not treat login success as safety assurance. Scope AI system access to the minimum required and review privileged access independently from model testing.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsAccess scoping and entitlement governance are the identity controls contrasted with behavioural safety testing.
Recommendation — Apply entitlement controls to AI applications separately from red-teaming and runtime safety checks.

Key terms

  • Behavioral Red-Teaming: Adversarial testing designed to make an AI system fail by manipulating prompts, context or inputs. The goal is to expose unsafe outputs, misalignment or jailbreak conditions before production use. It measures model behaviour, not user identity, access scope or entitlement quality.
  • Enterprise Authentication Stack: The set of identity components an application uses to authenticate users, federate logins, and manage access at scale. In practice, it combines application login, directory integration, session handling, and administrative controls so the application can support enterprise requirements without ad hoc workarounds.
  • AI Safety: AI safety is the discipline of preventing an AI system from taking unintended or harmful actions on its own. It focuses on the behaviour the system generates, even when no external attacker is involved. For identity teams, safety is about limiting what the agent can do once it is already operating.
  • Access Certification: Access certification is the periodic review of whether an identity still needs its current entitlements. For NHIs, certification is only reliable when reviewers know the identity's owner, purpose, and expiry, otherwise stale machine access can persist long after the original use case has ended.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org