By NHI Mgmt Group Editorial TeamDomain: AI SecuritySource: Holistic AIPublished November 13, 2025

TL;DR: Red teaming of open-source and open-weight Chinese models found strong performance alongside uneven safety, with jailbreak resistance ranging from 32% to 100% across tested models and safe-response rates spanning 81% to above 99%, according to Holistic AI. The result is a governance problem, not a model-quality problem: enterprises need independent testing, runtime guardrails, and ongoing observability before production use.


At a glance

What this is: This is Holistic AI’s red-teaming analysis of recent open Chinese generative AI models, and its key finding is that performance gains do not translate into consistent safety or jailbreak resilience.

Why it matters: It matters because AI governance teams need to validate model behaviour, not assume vendor claims hold under prompt injection, role-play, or policy-sensitive use cases, especially when models are being connected to enterprise data and workflows.

By the numbers:

👉 Read Holistic AI's analysis of red teaming open Chinese generative AI models


Context

Generative AI governance is failing most clearly at the point where model capability meets adversarial testing. Open and open-weight models can be attractive because they lower deployment barriers, but that does not mean they are safe to connect to enterprise workflows, especially when prompt injection, role-play abuse, and policy-sensitive prompts are in scope. This article is about that gap in AI governance, not just about which models perform well.

For AI security teams, the real issue is whether safety controls are validated independently and continuously rather than inferred from benchmark performance. The article shows why red teaming, runtime guardrails, and observability matter once models begin handling internal content, user input, or tool-using workflows. That starting position is increasingly common across enterprise AI programmes, not an edge case.


Key questions

Q: How should security teams evaluate GenAI models before production?

A: Security teams should test models with realistic adversarial scenarios, including direct prompt attacks and indirect instruction injection through retrieved content. The goal is to measure whether the model maintains its intended behavior under pressure. Approval should depend on repeatable evidence, not on a one-time benchmark score or vendor assurance.

Q: Why do open-weight models create extra governance risk for enterprises?

A: Open-weight models often shift more responsibility onto the organisation for testing, guardrails, and monitoring. Local deployment can improve privacy and control, but it also removes the assumption that the provider is handling safety end to end. That makes internal assurance, continuous monitoring, and access governance central to safe use.

Q: What do security teams get wrong about jailbreak testing?

A: They often treat jailbreak resistance as a one-time benchmark result instead of a control that can degrade with prompts, integrations, and user behaviour. A model that resists one class of attack may still fail under role-play, context shifts, or multilingual abuse. Testing has to reflect actual enterprise usage, not a lab-only scenario.

Q: What should organisations do when agentic AI starts using enterprise tools?

A: Organisations should define what the system may access, what actions require approval, and who is accountable if behaviour changes during execution. The key is to govern runtime authority, not just initial provisioning. Without that boundary, the AI workflow can expand its own operational reach faster than conventional IGA can observe it.


Technical breakdown

Why benchmark performance does not equal trustworthy behaviour

Model performance and model trustworthiness are related but not interchangeable. A system can answer questions well while still failing under adversarial prompts, context shifts, or policy-sensitive scenarios. Red teaming exposes whether a model preserves refusal behaviour, avoids unsafe completions, and resists prompt manipulation when the input is designed to bypass guardrails. For open-weight models, this matters even more because local deployment often shifts responsibility for control design to the organisation, not the model provider. Practical implication: validate models with adversarial test sets before production, not after users begin discovering failure modes.

Practical implication: require pre-production red teaming against harmful prompts, role-play attacks, and policy exceptions before any enterprise rollout.

How jailbreak resilience works in generative AI systems

Jailbreak resilience is the ability of a model and its surrounding controls to resist attempts to override policy through persuasion, instruction hierarchy abuse, or contextual manipulation. In practice, attackers do not need to break cryptography or exploit infrastructure if they can coerce the model into ignoring safety constraints. This is why role-play prompts, fictional scenarios, and language-shift attacks are so useful to adversaries. The model may still appear usable in normal interactions while becoming unreliable under structured abuse. Practical implication: treat jailbreak testing as a control validation exercise, not a model quality score.

Practical implication: assess jailbreak resilience as a control property and track it separately from general benchmark performance.

Why runtime guardrails and observability are now governance controls

Runtime guardrails are the policy layer that intercepts, filters, or transforms model outputs during live use, while observability tracks whether safety performance degrades over time. These controls matter because model behaviour can drift as prompts, users, context, and integrations change. In enterprise settings, the identity and access layer also becomes relevant when models interact with tools, repositories, or internal data sources, because unsafe output can quickly become unsafe action. Governance therefore spans both content control and access control. Practical implication: pair output filtering with telemetry that can prove when a model started to deviate.

Practical implication: implement live monitoring, content guardrails, and audit trails for any model connected to enterprise tools or sensitive data.


Threat narrative

Attacker objective: The attacker aims to coerce the model into producing unsafe output or bypassing policy so that enterprise workflows inherit the model's failure.

  1. Entry occurs through prompt injection, role-play, or context-shift prompts that target model instruction hierarchy rather than infrastructure.
  2. Escalation happens when the model bypasses its own refusal patterns and produces unsafe, policy-violating, or operationally risky output.
  3. Impact follows when that output is used in downstream workflows, creating compliance exposure, unsafe decisions, or data-handling failures.

NHI Mgmt Group analysis

Open-source model governance now depends on independent control validation, not adoption assumptions. The article shows that capability improvements do not eliminate safety variance across models, especially when adversarial prompting is involved. That means AI governance cannot treat openness, cost, or local deployment as proxies for trustworthiness. Organisations need evidence from red teaming, runtime testing, and policy enforcement before they scale usage. The practitioner conclusion is simple: model choice is only half the decision; control validation is the other half.

Prompt-injection resilience is becoming a named governance gap in enterprise AI programmes. The most useful concept from this article is the safety-control delta, the gap between a model's benchmark performance and its behaviour under coercive prompts. This gap matters because procurement teams often evaluate capability first and assume safety will generalise. It does not. The field should treat safety degradation under attack as a measurable control failure, not a theoretical concern. Practitioners should build governance around how models fail, not only how they perform.

AI governance is moving from static review to continuous assurance. The article’s emphasis on dynamic filtering and observability reflects a broader reality: AI systems are not one-time deployments. Their risk profile shifts as prompts, tools, data, and users change, which makes periodic approvals insufficient. For identity and access teams, the intersection is especially important when models can call tools or access enterprise data, because unsafe model output can turn into risky action. The practitioner conclusion is to align AI governance with continuous control monitoring.

Agentic workflows raise the stakes because unsafe model output can become unsafe execution. When models are integrated into tool-using systems, the control problem is no longer limited to text quality. Model outputs can drive actions, permissions requests, or automated next steps, which makes content safety part of operational security. That is where AI governance intersects with IAM and NHI controls: tool access, delegated permissions, and auditability all become part of the threat surface. The practitioner conclusion is to govern model-to-tool delegation as tightly as model output.

Production readiness should be judged by containment, not by headline performance metrics. A model that scores well on general tasks but fails on jailbreak resistance can still create enterprise risk once it is embedded in customer support, coding, or workflow automation. The right question is whether the organisation can contain bad outputs, explain them, and stop their propagation. This shifts evaluation away from marketing claims and toward measurable operational controls. The practitioner conclusion is to require containment evidence before any model reaches production.

What this signals

The practical signal for AI programmes is that benchmark wins do not reduce governance load. As soon as models are connected to tools, content filters, audit trails, and approval boundaries become mandatory control layers, not optional extras.

Safety-control delta: the gap between benchmark performance and adversarial behaviour is now a programme risk metric. Teams that cannot measure it will struggle to explain why a model that looked safe in testing failed in production.

Identity and access teams should pay particular attention when models are allowed to call tools, retrieve records, or trigger actions. That is where AI governance intersects with delegated privilege, and where auditability needs to extend beyond prompts to the actions the model can initiate.


For practitioners

  • Run adversarial red teams before production Test models against role-play, multilingual abuse, prompt injection, and context-shift scenarios before they are approved for business use. Keep the test set separate from the vendor benchmark so you can see how the model behaves in your environment.
  • Separate capability scoring from safety scoring Track utility, refusal behaviour, and jailbreak resilience as distinct metrics so a high benchmark score cannot mask weak containment. Use different thresholds for experimentation, pilot, and production.
  • Deploy runtime guardrails around sensitive prompts Interpose policy-based filters for unsafe content, tool requests, and data-bearing prompts before the model reaches downstream systems. The control should block or rewrite outputs that would otherwise create compliance or security exposure.
  • Add continuous observability to model behaviour Monitor prompt patterns, refusal drift, and anomalous outputs over time so control degradation is visible before it becomes a programme-wide issue. Feed the telemetry into audit and incident response workflows.
  • Govern tool use as a delegated access problem When models can call tools or access enterprise data, treat that delegation as a privileged workflow with logging, approval boundaries, and revocation paths. This is where model governance intersects with identity and access control.

Key takeaways

  • The article shows that performance and safety are not the same thing in generative AI.
  • Red teaming, runtime guardrails, and observability are the controls that turn model evaluation into enterprise governance.
  • When models can call tools, AI safety becomes an identity and access issue as much as a content issue.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI RMF, NIST AI 600-1 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST AI RMFMEASUREThe article centres on measuring model safety under adversarial prompts.
OWASP Agentic AI Top 10Agentic and tool-using AI raises prompt-injection and unsafe-output concerns.
NIST AI 600-1Generative AI governance and testing are directly relevant to these models.
NIST CSF 2.0PR.DS-1Runtime guardrails and observability support data and output protection.
MITRE ATLASTA0005 , Defense Evasion; TA0006 , Credential AccessPrompt abuse and unsafe delegation align with adversarial AI tactics.

Apply the GenAI profile to require pre-deployment testing, monitoring, and incident-ready documentation.


Key terms

  • Safe-response rate: The proportion of model outputs that remain aligned with policy when tested against harmful, borderline, or sensitive prompts. It is a governance metric, not just a benchmark score, because it shows how often the model preserves intended safety behaviour under pressure.
  • Jailbreak resilience: The ability of a model to resist attempts to override its safeguards through prompt injection, role-play, or contextual manipulation. It measures whether the model can keep its policy boundaries intact when an attacker tries to persuade it to ignore them.
  • Runtime Guardrail: A control applied while an AI agent is operating, not just during configuration or review. Guardrails can block dangerous tool calls, require approval for sensitive actions, or stop data leakage before it reaches systems or users.
  • Safety-control delta: The gap between how a model appears in ordinary benchmark testing and how it behaves under adversarial or enterprise-specific prompts. This gap matters because production risk is created by the difference between expected and observed behaviour, not by benchmark scores alone.

What's in the full article

Holistic AI's full blog covers the operational detail this post intentionally leaves for the source:

  • The per-model red team methodology, including how safe-response rate and jailbreak resilience were measured across the prompt set.
  • The full benchmark table showing which models passed or failed specific adversarial scenarios, useful for procurement and risk review.
  • The platform workflow for automated red teaming, runtime guardrails, and continuous observability in production AI environments.
  • The audit-reporting outputs that teams can use to document model behaviour for internal governance and compliance review.

👉 The full Holistic AI post covers model-by-model safety results, jailbreak testing detail, and governance controls.

Deepen your knowledge

The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, agentic AI identity, and secrets management. It is designed for practitioners who need to connect identity controls to emerging AI risk.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 18, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org