Treat safety as a question of whether the system’s intended behaviour is acceptable, and security as a question of whether an attacker can redirect that behaviour. Give each risk class its own owner, control set, and production gate. If one review process covers both, accountability becomes vague and response decisions slow down.
Why the distinction matters in practice
Security teams should treat AI safety and AI security as related but different control problems. Safety asks whether the system is behaving as intended in a way the organisation can accept; security asks whether someone can deliberately steer that behaviour. If those concerns are blended into one review, teams tend to miss the different owners, different evidence, and different escalation paths each class needs.
That separation becomes especially important when an AI system is both useful and externally reachable. A model can be safe enough for ordinary use and still be vulnerable to prompt injection, tool abuse, data exfiltration, or unsafe access to downstream systems. Conversely, a system can resist attack well and still produce outputs or actions that are operationally unsafe for the business.
For teams building governance around NIST AI RMF, the practical lesson is to avoid collapsing safety and security into one approval gate. Safety evidence should focus on acceptable model behaviour, while security evidence should focus on adversarial misuse, access boundaries, and containment.
How to split ownership, controls, and release gates
The cleanest operating model is to assign one owner for intended behaviour and a separate owner for adversarial misuse. That does not mean separate teams must never talk to each other, but it does mean the decisions are different: safety owners judge policy, content, and outcome acceptability; security owners judge trust boundaries, access paths, and abuse cases.
A useful rule is that safety reviews answer, “Would we still ship this if nobody attacked it?” Security reviews answer, “Could an attacker redirect the system or its tools to cause harm?” When the answer to either is no, the issue should block release for its own reason, not be absorbed into a generic AI sign-off.
Where the system uses tools, APIs, or agentic actions, the security gate should explicitly cover identity and privilege abuse, because a benign model can still become dangerous if an attacker redirects its authority. Safety testing alone will not surface that class of failure.
What teams should test separately before production
Safety testing should prove the system remains within acceptable behavioural bounds under normal and foreseeable use. That includes harmful output tolerance, policy compliance, refusal quality, and whether the system produces actions the business would regard as unacceptable even when no attacker is present.
Security testing should prove the system resists manipulation by hostile inputs, compromised integrations, or stolen credentials. For agentic systems, the focus should extend to tool permissions, session boundaries, secrets handling, and whether external content can change what the system does in a way the owner did not intend.
For AI systems that rely on APIs or external services, API security controls help teams test whether the model or agent can overreach into functions it should not call. That gives the security review a concrete control surface instead of a vague discussion about “model risk.”
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP API Security Top 10 address the attack surface, NIST AI RMF and NIST SP 800-53 Rev 5 set the technical controls, and ISO/IEC 42001:2023 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | Govern | AI safety and security separation is an AI risk governance issue. |
| Recommendation — Split AI safety and security decisions into separate governance controls and owners. | ||
| OWASP Agentic AI Top 10 | ASI03 — Identity & Privilege Abuse | Security failure in agentic AI often comes from redirected authority and privilege. |
| Recommendation — Review tool access and privileges separately from behaviour safety before release. | ||
| OWASP API Security Top 10 | API5 — Broken Function Level Authorization | AI systems with tools and APIs can fail securely only if function access is bounded. |
| Recommendation — Enforce function-level authorization for model and agent actions. | ||
| NIST SP 800-53 Rev 5 | IA-9 — Identification and Authentication (Service and User Accounts) | AI systems often rely on service or machine access that must be governed separately from safety. |
| Recommendation — Authenticate service accounts and constrain their access paths independently of model safety checks. | ||
| ISO/IEC 42001:2023 | A.5.2 — AI policy | AI programmes need policy separation between acceptable behaviour and security assurance. |
| Recommendation — Define separate policy gates for safety review and security review. | ||
Practitioner Guidance
What to prioritise: Separate the review artefacts first, then decide where overlap is genuinely unavoidable. If the same committee signs off both safety and security, require distinct checklists, distinct approvers, and distinct go or no-go criteria.
What to verify: Confirm that the safety gate is judging acceptable behaviour, while the security gate is judging adversarial control, tool reach, and privilege boundaries. If you cannot point to different failure modes for each gate, the operating model is still too vague.
Common mistake: Treating “the model behaved badly” and “the model was attacked” as the same issue. That shortcut usually delays response, because the remediation path for unsafe behaviour is not the same as the remediation path for compromised behaviour.
Practitioner takeaway: The goal is not to create more process, but to make accountability concrete: one owner for behavioural acceptability, one owner for attack resistance, and one release decision that cannot hide behind blended language.