Join our Newsletter — 33% off our NHI Course
Home FAQ AI Security How should security teams govern GenAI safety when…
AI Security

How should security teams govern GenAI safety when Trust and Safety and AI teams are split?

← Back to all FAQ
By NHI Mgmt Group Editorial Team Updated August 21, 2026 Domain: AI Security

They should assign explicit ownership for policy, abuse taxonomy, red teaming, monitoring, and enforcement before launch. When those responsibilities are split across separate teams, gaps appear in escalation and accountability. The most effective model is shared governance with clear control handoffs, so model work, moderation, and identity-based enforcement operate as one programme.

Why This Matters for Security Teams

When Trust and Safety and AI engineering operate as separate functions, GenAI risk often lands in the gap between policy intent and technical enforcement. That matters because safety failures are not limited to harmful outputs. They also include prompt injection, unsafe tool use, data leakage, policy drift, and weak escalation paths when abuse patterns change faster than the review process. The right operating model is less about who “owns AI” and more about who can enforce controls across the full lifecycle, from model change to moderation action. The NIST Cybersecurity Framework 2.0 is useful here because it forces teams to think in terms of governance, protection, detection, response, and recovery rather than isolated functions.

The common mistake is treating safety as a content review problem after launch instead of a design and operational control problem before launch. If policy authors, abuse analysts, red teamers, and platform engineers do not share a single control model, then enforcement becomes inconsistent and incident response slows down. In practice, many security teams encounter this only after a harmful prompt, model leak, or moderation miss has already reached production users, rather than through intentional cross-functional control design.

How It Works in Practice

Effective governance starts by splitting responsibilities without splitting accountability. Trust and Safety can define the abuse taxonomy, policy thresholds, escalation criteria, and user-facing enforcement actions. AI teams can own model behaviour, evaluation harnesses, guardrail tuning, and release gating. Security teams should coordinate logging, detection, incident response, and identity-based enforcement where accounts, API keys, or privileged operator access are part of the risk surface. The goal is a shared operating model with named control owners and documented handoffs.

A practical implementation usually includes:

  • One policy baseline for unacceptable content, unsafe assistance, tool abuse, and disallowed data handling.
  • Pre-launch red teaming that tests prompt injection, jailbreaks, exfiltration attempts, and unsafe autonomous actions.
  • Runtime monitoring for abuse trends, policy exceptions, and anomalous tool calls or high-risk prompts.
  • Enforcement paths that connect moderation decisions to account actions, rate limits, workflow approval, or credential revocation.
  • Change control for model updates, retrieval sources, tools, and system prompts so safety regressions are reviewed before release.

The NIST AI 600-1 GenAI Profile is especially relevant because it translates AI risk management into operational expectations for mapping risks, measuring them, and documenting controls. For teams with agentic features, current guidance suggests extending the same model to tool permissions and identity governance, since an AI system with execution authority changes the meaning of “safety” from output quality to action control. These controls tend to break down when model deployments are frequent and moderation workflows remain manual, because the safety decision path cannot keep pace with the release cadence.

Common Variations and Edge Cases

Tighter safety governance often increases coordination overhead, requiring organisations to balance speed of experimentation against control consistency. That tradeoff becomes more visible when teams are decentralised, because local product groups may want different thresholds for refusal, escalation, or age-sensitive content. There is no universal standard for this yet, but best practice is evolving toward a central policy authority with delegated implementation.

Some environments need additional nuance. Consumer-facing chat products often focus on abuse prevention and harmful content escalation. Enterprise copilots may instead prioritise data loss prevention, prompt injection resistance, and role-based access. Agentic systems raise a further issue: if the model can call tools, send messages, or trigger workflows, then Trust and Safety cannot operate as a post-hoc moderation layer alone. Identity controls, approval gates, and scoped permissions become part of the safety design. Security teams should also treat vendor models and hosted tools as supply chain dependencies, because the governance burden does not end at the application boundary.

The practical question is not whether Trust and Safety or AI owns the problem, but whether both functions can act on the same control evidence when abuse patterns change. Shared governance fails when teams share a dashboard but not decision rights, especially in high-change environments with multiple model versions, retrieval sources, and external tool integrations.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATLAS and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST AI RMF, NIST AI 600-1 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST AI RMFAI risk governance fits split ownership and control handoffs for GenAI safety.
MITRE ATLASATLAS covers adversarial tactics like prompt injection and model abuse patterns.
OWASP Agentic AI Top 10Agentic features add tool abuse and execution risks that need explicit safety controls.
NIST AI 600-1GenAI profile guidance helps translate safety risks into operational control expectations.
NIST CSF 2.0GV.OV-01Shared governance needs clear oversight and accountability across teams.

Use AI RMF to define governance, measurement, and monitoring across the shared safety model.

NHIMG Editorial Note
Reviewed and updated by the NHIMG editorial team on August 21, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org