TL;DR: DeepSeek-R1 scores higher risk than o3-mini across EU AI Act compliance, privacy, security, fairness, and adversarial robustness, according to VirtueAI’s comparative red-teaming analysis, while both models still need guardrails before broad deployment. The result is a governance problem, not just a model-quality issue: AI safety controls must now track regulatory exposure, data leakage, and misuse pathways together.
At a glance
What this is: VirtueAI compares OpenAI o3-mini and DeepSeek-R1 and concludes that DeepSeek-R1 presents materially higher safety and compliance risk across several evaluation dimensions.
Why it matters: For IAM and AI governance teams, the finding matters because model risk, privacy leakage, and access to sensitive data now intersect with identity, control ownership, and deployment approval.
👉 Read VirtueAI's comparative red-teaming analysis of o3-mini and DeepSeek-R1
Context
AI model red-teaming is the practice of probing a model for unsafe behaviour, including privacy leakage, harmful outputs, prompt injection susceptibility, and policy bypass. In this case, the core governance gap is that model safety can no longer be treated as a single score, because compliance risk, data leakage, and adversarial robustness fail differently and need different controls.
That matters to AI governance, IAM, and security teams because foundation models increasingly sit inside business workflows that touch personal data, regulated decisions, and connected tools. When a model is permitted to generate, classify, or recommend actions, the control problem extends beyond the model itself to the identities, permissions, and data paths around it.
Key questions
Q: How should organisations approve AI models for real-world use?
A: Approve models by use case, not by headline benchmark alone. Define acceptable failure modes for privacy, bias, deception, and adversarial robustness, then require evidence that the model stays within those limits in the exact workflow it will support. Production approval should also include data classification, human oversight, and a rollback path if the model drifts.
Q: Why do AI models create governance risk even without retraining?
A: Because behaviour can change at inference time when the model sees new context, examples, or instructions. That means access decisions made before a session starts are not enough on their own. Practitioners need controls that address what the model can consume and do during execution, not only what it was permitted to access originally.
Q: What breaks when an AI tester has broad tool access?
A: Broad tool access makes the agent harder to audit, easier to misdirect, and more likely to overreach its intended scope. It also increases the chance that credentials, findings, or commands are reused in the wrong context, which turns a helpful assistant into an uncontrolled operator.
Q: Who is accountable when AI output causes a compliance or legal issue?
A: Accountability sits with the organisation that deploys and governs the AI use case, not only with the vendor that hosts the model. If an employee or agent uses AI in a business context, the enterprise must be able to show policy, monitoring, and evidence of control. That is now a governance obligation, not optional hygiene.
Technical breakdown
How comparative red-teaming translates model behaviour into governance risk
Comparative red-teaming evaluates two or more models against the same adversarial and policy-based prompts so evaluators can see where behaviour diverges. The useful output is not just pass or fail, but a matrix of failure modes across safety, privacy, fairness, and robustness. That matters because a model can be acceptable for one use case and unacceptable for another. For governance, the key is to map each failure mode to a deployment boundary, approval condition, or compensating control rather than treating the benchmark as a universal safety verdict.
Practical implication: tie model approval to use-case-specific thresholds, not to a single overall safety label.
EU AI Act and GDPR risks in model evaluation
The article frames model assessment through the EU AI Act and GDPR, which is important because AI systems that process personal data or influence consequential decisions create compliance exposure as well as technical risk. In practice, privacy leakage, deceptive outputs, and biased automated decisions can trigger governance obligations around transparency, data minimisation, accountability, and human oversight. For teams, this means safety evaluation must be documented as part of the control environment, not left as an informal research activity. The governance question is whether the model can be justified in the specific process where it will operate.
Practical implication: connect model testing evidence to the compliance review path before any production rollout.
Why adversarial robustness is an identity and access problem as well as an AI problem
Adversarial robustness is about how well a model resists jailbreaks, prompt injection, deceptive instruction chains, and other manipulations that change its behaviour. In connected environments, those attacks often become an access problem because the model may be granted tool access, retrieval access, or delegated workflow actions. If the model can be induced to reveal secrets or call internal systems, the failure is not only unsafe content generation but over-broad privilege. That is where AI governance meets IAM and NHI controls: the model’s permissions, session scope, and data reach must be explicitly bounded.
Practical implication: treat model tool access like privileged access and scope it to the minimum task needed.
NHI Mgmt Group analysis
Model safety scoring is becoming a governance control, not a research metric. VirtueAI’s comparison shows that model evaluation now influences deployment approval, legal exposure, and operational trust. A model that scores poorly on privacy or deceptive output cannot be treated as merely less accurate, because the downstream business risk is materially different. Practitioners should treat red-team evidence as part of the control record, not as supplementary commentary.
AI governance debt is accumulating where organisations deploy models faster than they can define acceptable failure modes. The article’s split across compliance, safety, fairness, and robustness shows why a single pass or fail label is no longer enough. Teams need explicit decision criteria for each use case, especially where models touch personal data or regulated workflows. The practitioner conclusion is simple: if you cannot name the failure mode you are willing to accept, you cannot govern the deployment responsibly.
Identity governance now extends to model-enabled actions, not just human or service accounts. When a model is allowed to call tools, generate decisions, or access data stores, it becomes part of the trust boundary. That requires least privilege, bounded delegation, and auditable action paths. For identity programmes, the lesson is to bring AI systems into the same control language used for NHI and privileged access.
DeepSeek-R1’s higher risk profile reinforces the need for deployment gates, not blanket enablement. The article indicates that one model can be materially weaker across privacy and adversarial resilience even when headline performance looks similar. That makes contextual approval essential. Practitioners should move from model enthusiasm to controlled rollout decisions tied to data sensitivity and business criticality.
What this signals
Model governance will increasingly look like access governance. As AI systems gain tool use, retrieval, and workflow actions, security teams will need explicit permission models, session boundaries, and evidence that each model action is traceable to an approved use case. The practical shift is from evaluating outputs to governing delegated behaviour.
AI governance debt: the backlog created when organisations deploy models faster than they define failure modes, approval thresholds, and rollback criteria. Teams that do not close that gap will keep expanding AI use without knowing where risk actually sits.
For identity and security programmes, the next control question is not whether the model is capable, but what it is authorised to do. That means aligning model access with privileged access principles, then testing whether the model can be constrained as tightly as a human administrator or service account.
For practitioners
- Define use-case-specific model approval gates Separate evaluation thresholds for regulated decision support, public-facing chat, and internal knowledge workflows. Require evidence for privacy, deception resistance, and adversarial robustness before production access is granted.
- Map model failures to compliance obligations Document how privacy leakage, bias, and unsafe decision-making affect EU AI Act and GDPR obligations in each workflow. Keep the approval record linked to the actual data categories and business decisions involved.
- Constrain model tool and data access Treat any model that can call tools or retrieve enterprise data as a privileged actor. Scope retrieval, action permissions, and session duration to the minimum needed for the task.
- Require repeatable red-team evidence Run the same prompt suites across model versions and keep the results in change control. Compare failure patterns over time so security, legal, and product teams can see whether risk is improving or drifting.
Key takeaways
- VirtueAI’s comparison shows that model risk is multidimensional, with compliance, privacy, bias, and adversarial resilience needing separate evaluation.
- A model that performs well on factual tasks can still create governance exposure if it leaks data, resists poorly to manipulation, or influences regulated decisions.
- Practitioners should govern AI models with the same discipline used for privileged access, including scoped permissions, documented approval criteria, and repeatable red-team evidence.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST AI RMF and NIST CSF 2.0 set the technical controls, while EU AI Act and GDPR define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | GOVERN | The post is about AI risk governance and accountable oversight of model deployment. |
| EU AI Act | Art.9 | The article directly evaluates compliance risk under the EU AI Act. |
| GDPR | Art.32 | Privacy leakage and personal-data handling create GDPR security and protection obligations. |
| NIST CSF 2.0 | GV.RM-01 | The article is about managing AI risk as part of enterprise governance. |
Assess whether model access, outputs, and logging meet security of processing requirements.
Key terms
- Red Teaming: Red teaming is structured adversarial testing used to find how an AI system fails under realistic misuse or attack conditions. In AI security, it is a discovery method, not a proof of safety, because probabilistic behaviour and changing models prevent any lasting guarantee.
- Adversarial Robustness: Adversarial robustness is a model’s ability to behave safely when inputs are manipulated, unusual, or intentionally crafted to cause failure. In practice, it is measured through testing and red-teaming, not assumed from functional accuracy, and it becomes a core control when AI systems move into production.
- Model Governance: Model governance is the set of controls that decides which foundation models can be used for which agent types and use cases. It links platform choice to security policy, because the model selection influences data exposure, tool behaviour, and the risk profile of the resulting agent.
- Automated Decisioning: Automated decisioning is the use of software or models to make or trigger business actions without manual approval for each case. It increases speed and scale, but it also shifts control away from human review and toward the quality of the underlying logic, data, and auditability.
What's in the full article
VirtueAI's full post covers the operational detail this post intentionally leaves for the source:
- Model-by-model evaluation notes for o3-mini and DeepSeek-R1 across safety, privacy, fairness, and robustness dimensions.
- Examples of deceptive output, hallucination, and policy-violating behaviour that are useful for hands-on AI risk review.
- The red-teaming framing behind the EU AI Act and GDPR risk assessments used in the comparison.
- Illustrative outputs showing how the models behave under automated decision-making and fraudulent prompt scenarios.
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, identity lifecycle, and secrets management. It helps security and identity practitioners apply control thinking to human, machine, and agentic access patterns.
Published by the NHIMG editorial team on August 18, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org