TL;DR: As GenAI safety responsibility shifts from Trust and Safety into Responsible AI and AI safety teams, ActiveFence’s Alice argues that T&S leaders risk losing budget, influence, and practical safety coverage unless they partner early, shape policy, and own post-launch monitoring. The underlying governance problem is not just organisational politics, but a widening gap between model testing and the real-world abuse patterns that emerge after deployment.
At a glance
What this is: This is an analysis of how GenAI safety responsibilities are moving away from Trust and Safety teams and why that creates governance gaps across model policy, abuse mapping, and post-launch monitoring.
Why it matters: It matters because identity, moderation, and abuse controls are now part of AI governance, and teams responsible for IAM, identity verification, and agent oversight need to coordinate with safety functions before gaps turn into operational risk.
By the numbers:
- 92% agree governing AI agents is critical to enterprise security, yet only 44% have implemented any policies to do so.
- 72% of organisations have experienced or suspect they have experienced a breach of non-human identities, 46% confirmed and 26% suspected.
👉 Read ActiveFence's analysis of GenAI safety, Trust and Safety, and AI governance
Context
GenAI safety programmes often fail when they are treated as a model-only concern rather than a broader operating model problem. In practice, the security gap appears when teams responsible for moderation, abuse review, identity verification, and policy enforcement are excluded from AI safety decisions, even though their controls are the ones that shape what happens after deployment.
For identity and security practitioners, the real issue is governance ownership. When AI safety, Trust and Safety, and Responsible AI teams are split across budgets and reporting lines, the programme can lose the people who understand user abuse, credential misuse, behavioural escalation, and post-launch monitoring. That split is increasingly common as organisations expand GenAI adoption.
The article’s starting position is typical of large technology organisations that are reorganising around GenAI, not an edge case.
Key questions
Q: How should security teams govern GenAI safety when Trust and Safety and AI teams are split?
A: They should assign explicit ownership for policy, abuse taxonomy, red teaming, monitoring, and enforcement before launch. When those responsibilities are split across separate teams, gaps appear in escalation and accountability. The most effective model is shared governance with clear control handoffs, so model work, moderation, and identity-based enforcement operate as one programme.
Q: Why do GenAI safety programmes need identity and user-behaviour controls?
A: Because many real-world failures happen after launch, when harmful users, repeated abuse patterns, and evasive behaviour emerge in production. Identity and behavioural controls let teams score accounts, investigate abuse, and contain repeat offenders. Without those controls, safety efforts stay limited to model quality and miss the operational layer where abuse scales.
Q: What do organisations get wrong about AI safety and access control?
A: Organisations often focus on model outputs while ignoring the privileges behind the model. If an agent can read sensitive data or invoke tools, the real risk is what it can cause the environment to do. Effective control starts with scope, policy, and monitoring around actions, not just moderation of generated text.
Q: Who should own accountability for deployed AI agents?
A: Accountability should sit with the business or governance owner who can approve scope, review changes and retire the agent when it is no longer needed. Shared ownership without clear decision rights usually turns into no ownership, which is how agents become difficult to audit and even harder to decommission.
Technical breakdown
How GenAI safety programmes split between model and system controls
GenAI safety is often divided into pre-launch model controls and post-launch system controls. Model controls focus on training data, evaluation sets, red teaming, and output filtering, while system controls cover user behaviour, moderation workflows, flagging, investigations, and policy enforcement in live environments. That split matters because model testing cannot predict every abuse path once a product is exposed to real users. The practical failure is assuming a safe model automatically creates a safe service. In reality, the service boundary includes humans, prompts, workflows, and escalation paths that change the threat surface after release.
Practical implication: Practitioners should map safety ownership across the full lifecycle, not only the model build phase.
Why abuse taxonomy and policy design are governance controls
Risk mapping is not just documentation. In GenAI environments, the abuse taxonomy defines what the system should detect, refuse, escalate, or block, and policy design turns that taxonomy into operational controls. If policy work is left to teams that understand fairness but not real abuse patterns, the organisation can miss categories such as self-harm, child exploitation, misinformation, or coordinated harmful behaviour. The result is a policy that looks complete on paper but does not cover the ways adversaries actually use the system. This is a governance failure as much as a safety failure.
Practical implication: Security and Trust and Safety teams should co-author abuse taxonomies and policy exceptions before launch.
Why identity verification matters in platform-level AI safety
Identity verification becomes relevant when a platform needs to distinguish legitimate users from repeat abusers, high-risk accounts, or coordinated actors. In GenAI operations, identity signals can support user scoring, friction, investigation, and access limitation, especially where abuse is persistent rather than one-off. That does not mean every system needs full real-name verification. It means AI safety teams need a defensible way to connect behaviour to identity signals when the abuse pattern warrants escalation. Without that, moderation becomes reactive and attribution remains weak, which limits containment.
Practical implication: Teams should define where identity verification, account scoring, and enforcement belong in the abuse response workflow.
Threat narrative
Attacker objective: The objective is to exploit weak GenAI safety governance to produce harmful content, evade moderation, or scale abuse without effective containment.
- Entry occurs when threat actors or problematic users reach a GenAI platform that has weak abuse gating or incomplete policy coverage.
- Escalation follows when the system lacks post-launch monitoring, allowing repeated harmful prompts, evasions, or misuse patterns to continue unchecked.
- Impact is realised when the organisation loses control of moderation outcomes, safety coverage, or user trust, and the AI programme becomes harder to govern at scale.
NHI Mgmt Group analysis
Trust and Safety is becoming an identity-adjacent control layer in GenAI governance. The article shows that safety work is no longer confined to model quality or content moderation. Once platforms need to identify abusive users, score accounts, and connect behaviour to enforcement, Trust and Safety becomes part of the wider identity control plane. That is why IAM, identity verification, and access governance teams should treat AI safety as a shared operating model, not a separate specialist lane.
The named concept here is safety ownership drift. This is the pattern in which responsibility moves from the people who understand abuse and user behaviour to newly formed AI teams that may understand models better than operations. The drift creates policy gaps, budget fragmentation, and weaker escalation paths. Once that happens, the programme may still launch, but it launches with incomplete governance and limited real-world feedback loops.
GenAI safety programmes fail when they privilege model integrity over service integrity. The article correctly separates model testing from live system monitoring, and that distinction matters. A model can be evaluated well and still be unsafe in production if user behaviour, abusive workflows, or reporting mechanisms are not governed. The field needs to stop treating launch as the end of safety work.
Identity verification has a legitimate place in abuse containment, but only when tied to clear enforcement logic. The article’s point about user flagging and scoring is important because it shows where identity can support moderation without becoming a blunt privacy instrument. The practitioner lesson is that identity signals must be purpose-built for abuse response, not bolted on after incidents create pressure.
What this signals
The programme lesson is that safety teams cannot afford to treat moderation, abuse review, and identity verification as after-the-fact add-ons. As GenAI grows, the security boundary shifts toward runtime governance, and the organisation needs controls that survive both user pressure and product velocity.
Safety ownership drift: when AI safety moves away from the teams that understand abuse behaviour, the control environment weakens even if the model itself is well tested. Practitioners should watch for this drift whenever budgets, reporting lines, or tooling ownership change, because it usually precedes fragmented escalation and weaker containment.
Identity and governance teams should also prepare for more explicit links between behavioural enforcement and account assurance. That means aligning moderation workflows with identity verification policy, logging, and escalation evidence, then mapping those controls to NIST AI Risk Management Framework expectations where AI governance is in scope.
For practitioners
- Map GenAI safety ownership across teams Document which team owns policy, abuse taxonomy, red teaming, post-launch monitoring, and escalation so gaps do not emerge between Responsible AI and Trust and Safety. Include clear handoffs for moderation and incident response.
- Co-author abuse taxonomies with safety and identity teams Build shared risk maps that include harmful content, evasion tactics, repeat-abuse patterns, and identity-linked enforcement triggers. Use these maps to decide what the system must block, score, or escalate.
- Define post-launch monitoring as a control requirement Treat user flagging, investigation workflows, and behavioural scoring as part of the release criteria for any GenAI feature, not as a later enhancement. That closes the gap between model testing and live misuse.
- Set clear boundaries for identity verification in abuse response Specify when identity verification is justified, what evidence triggers it, and how it supports enforcement without becoming universal friction. This prevents reactive policy drift after abuse events.
Key takeaways
- GenAI safety becomes fragile when Trust and Safety expertise is removed from policy, moderation, and post-launch monitoring.
- The governance gap is not just organisational. It is the loss of abuse knowledge, escalation discipline, and identity-linked enforcement in live environments.
- Practitioners should treat safety ownership, identity signals, and runtime monitoring as connected controls, not separate workstreams.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 address the attack surface, NIST AI RMF, NIST AI 600-1 and NIST CSF 2.0 set the technical controls, and GDPR define the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | GOVERN | The article is about AI governance ownership and accountability. |
| NIST AI 600-1 | GenAI safety, evaluation, and deployment controls are directly in scope. | |
| OWASP Agentic AI Top 10 | Agent and model misuse patterns map to agentic application risks. | |
| NIST CSF 2.0 | PR.AC-4 | Access and identity-linked enforcement are part of the control gap described. |
| GDPR | Art.32 | Identity verification and behavioural enforcement can implicate personal data handling. |
Assess whether moderation and verification data processing satisfies security and minimisation requirements.
Key terms
- GenAI Safety: GenAI safety is the set of controls that reduce harmful, unsafe, or policy-violating outcomes from generative AI systems. It includes model testing, prompt and output controls, abuse detection, monitoring, escalation, and governance across the full product lifecycle.
- Trust And Safety: Trust and safety is the combined discipline of preventing abuse, reducing harm, and preserving legitimate participation in a digital community. In identity programmes, it links verification, moderation, and lifecycle governance so account confidence and user experience are managed together.
- Post-Launch Monitoring: Post-launch monitoring is the continuous review of live system behaviour after a service is released. For GenAI, it covers user abuse, policy evasions, harmful outputs, and escalation workflows, because many safety failures only appear in production, not in testing.
- Abuse Taxonomy: An abuse taxonomy is a structured classification of harmful behaviours, content types, and attack patterns a system must address. In AI safety, it guides what the model should refuse, what moderators should escalate, and what identity or behavioural signals should trigger enforcement.
What's in the full article
ActiveFence's full blog covers the operational detail this post intentionally leaves for the source:
- The article’s step-by-step breakdown of how T&S teams can map GenAI responsibilities to internal stakeholders and decision owners.
- The practical examples of policy development, filtering, and training-data collaboration that the source uses to show where teams lose influence.
- The discussion of post-launch mitigation systems, including user flagging, scoring, and incident handling in live GenAI environments.
- The source’s perspective on how T&S teams can stay relevant as AI safety budgets and ownership shift across the organisation.
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, identity lifecycle, and secrets management for practitioners who need to connect access control to broader security operations. It gives identity and security teams a shared language for governing machine access, user behaviour, and the controls around them.
Published by the NHIMG editorial team on August 21, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org