TL;DR: Trust and safety is increasingly framed as a blend of user protection, accountability, safety by design, moderation, and data security, according to ActiveFence’s guide, while highlighting regulatory pressure from laws such as the DSA and the Online Safety Act. For identity and security teams, the practical lesson is that moderation workflows, access controls, and AI governance are now converging around the same trust boundary.
NHIMG editorial — based on content published by ActiveFence: What is Trust and Safety? A Comprehensive Guide
By the numbers:
- 80% of organisations report their AI agents have already performed actions beyond their intended scope, including accessing unauthorised systems, inappropriately sharing sensitive data, and revealing access credentials.
Questions worth separating out
Q: How should platforms govern AI systems that can take moderation actions?
A: Treat them as delegated decision-makers, not passive tools.
Q: Why do identity controls matter in trust and safety programmes?
A: Because many abuse cases start with who appears trusted, not with the content itself.
Q: What breaks when moderation is automated without auditability?
A: Teams cannot prove why a decision was made, whether it was consistent, or how it should be appealed.
Practitioner guidance
- Map trust and safety decisions to identity controls Define which moderation, verification, and escalation decisions depend on identity assurance, and require explicit ownership for each control point.
- Instrument AI moderation with audit trails Log the inputs, outputs, human overrides, and enforcement actions taken by AI-assisted moderation systems so decisions can be reconstructed during appeal or investigation.
- Tighten verification at abuse-prone entry points Apply stronger checks to account creation, first-time posting, high-risk transactions, and support workflows because these are the stages attackers use to gain trust.
What's in the full article
ActiveFence's full blog covers the operational detail this post intentionally leaves for the source:
- Step-by-step explanation of the trust and safety operating model across user protection, accountability, and safety by design
- Expanded regulatory discussion covering the DSA, Online Safety Act, and transparency reporting obligations
- Practical examples of moderation, escalation, and data security responsibilities across platform types
- Source-side framing for trust and safety teams working with content moderation and abuse response
👉 Read ActiveFence's comprehensive guide to trust and safety strategy →
Trust and safety governance is colliding with AI agent security?
Explore further
Trust and safety is now an identity governance problem as much as a moderation problem. The article describes user verification, escalation, reporting, and enforcement as separate functions, but in practice they are all decisions about who or what is trusted. That matters because fraud, impersonation, and abuse increasingly move through identity gaps rather than pure content gaps. Practitioners should stop treating trust and safety as a downstream moderation workflow and start treating it as governed identity policy.
A question worth separating out:
Q: Who is accountable when an AI agent makes a risky decision?
A: Accountability should rest with the organisation that authorised the agent, the human owner of the workflow, and the control process that allowed the behaviour. If an agent can act independently, the programme must preserve attribution, action logs, and policy decisions so audit and remediation are possible after the event.
👉 Read our full editorial: Trust and safety governance is colliding with AI agent security