TL;DR: DeepSeek-R1 scores higher risk than o3-mini across EU AI Act compliance, privacy, security, fairness, and adversarial robustness, according to VirtueAI’s comparative red-teaming analysis, while both models still need guardrails before broad deployment. The result is a governance problem, not just a model-quality issue: AI safety controls must now track regulatory exposure, data leakage, and misuse pathways together.
NHIMG editorial — based on content published by VirtueAI: How Safe Are OpenAI o3-mini and Deepseek-R1? A Comparative Red-Teaming Analysis
Questions worth separating out
Q: How should organisations approve AI models for real-world use?
A: Approve models by use case, not by headline benchmark alone.
Q: Why do AI models create governance risk even without retraining?
A: Because behaviour can change at inference time when the model sees new context, examples, or instructions.
Q: What breaks when an AI tester has broad tool access?
A: Broad tool access makes the agent harder to audit, easier to misdirect, and more likely to overreach its intended scope.
Practitioner guidance
- Define use-case-specific model approval gates Separate evaluation thresholds for regulated decision support, public-facing chat, and internal knowledge workflows.
- Map model failures to compliance obligations Document how privacy leakage, bias, and unsafe decision-making affect EU AI Act and GDPR obligations in each workflow.
- Constrain model tool and data access Treat any model that can call tools or retrieve enterprise data as a privileged actor.
What's in the full article
VirtueAI's full post covers the operational detail this post intentionally leaves for the source:
- Model-by-model evaluation notes for o3-mini and DeepSeek-R1 across safety, privacy, fairness, and robustness dimensions.
- Examples of deceptive output, hallucination, and policy-violating behaviour that are useful for hands-on AI risk review.
- The red-teaming framing behind the EU AI Act and GDPR risk assessments used in the comparison.
- Illustrative outputs showing how the models behave under automated decision-making and fraudulent prompt scenarios.
👉 Read VirtueAI's comparative red-teaming analysis of o3-mini and DeepSeek-R1 →
Deepseek-R1 safety gaps: what AI governance teams should act on?
Explore further
Model safety scoring is becoming a governance control, not a research metric. VirtueAI’s comparison shows that model evaluation now influences deployment approval, legal exposure, and operational trust. A model that scores poorly on privacy or deceptive output cannot be treated as merely less accurate, because the downstream business risk is materially different. Practitioners should treat red-team evidence as part of the control record, not as supplementary commentary.
A question worth separating out:
Q: Who is accountable when AI output causes a compliance or legal issue?
A: Accountability sits with the organisation that deploys and governs the AI use case, not only with the vendor that hosts the model. If an employee or agent uses AI in a business context, the enterprise must be able to show policy, monitoring, and evidence of control. That is now a governance obligation, not optional hygiene.
👉 Read our full editorial: AI model red-teaming shows deepseek-r1 raises higher safety risk