By NHI Mgmt Group Editorial TeamBased on Abnormal AI: “The Adversary's New Assistant: Weaponizing AI Chatbots” (June 26, 2026)

TL;DR: Legitimate AI platforms are being manipulated to generate convincing phishing lures, malicious scripts, and automated fraud at scale, according to Abnormal AI and GPAI. That shifts the problem from content quality to identity, trust, and abuse controls across human, machine, and agentic workflows.


At a glance

What this is: This on-demand webinar says adversaries are abusing legitimate AI platforms to produce phishing, malicious code, and fraud content at scale.

Why it matters: It matters because IAM, fraud, and security teams now have to govern trusted AI workflows as abuse surfaces, not just content producers.


Context

Adversarial AI in this context means attackers using legitimate AI systems to help generate harmful content, automate deception, or scale fraud. The key governance gap is that many organisations still treat AI tools as productivity software rather than as abuse-prone execution environments that can be repurposed through prompts, workflows, and integrations.

Abnormal AI and GPAI frame the problem around phishing and fraud at scale, with the attacker using well-known platforms as a force multiplier rather than building everything from scratch. That shifts the control conversation toward trust boundaries, workflow governance, and how identity is handled across human users, API-connected services, and increasingly automated AI-assisted processes.


Key questions

Q: How should security teams govern agentic AI in fraud detection?

A: Start by separating detection support from decision authority. Agentic AI can correlate signals quickly, but any action that changes customer treatment, case priority, or compliance status needs explicit policy, traceability, and a human escalation path. Governance should define what the system may recommend, what it may execute, and what always requires review.

Q: Why do traditional content filters struggle against adversarial AI abuse?

A: Because the attacker can use a legitimate model to produce polished, varied, and context-aware output that looks normal at the surface. The real signal shifts from language quality to provenance, session behaviour, and whether the identity driving the workflow matches its expected purpose.

Q: What are the warning signs that an AI account is being used for abuse?

A: Look for repeated prompt refinement, unusual bursts of generation, output reused across campaigns, and access patterns that do not fit the account’s ordinary role. Those signals suggest the workflow is being steered toward phishing or fraud rather than legitimate business use.

Q: Should fraud teams and IAM teams respond separately to AI-enabled deception?

A: No. AI-enabled phishing and fraud sit at the intersection of identity governance, abuse detection, and campaign response. IAM must control who can use the tools and under what conditions, while fraud teams must absorb the new content and volume characteristics that AI introduces.


Background and context

How adversaries turn legitimate AI platforms into abuse infrastructure

The technical issue is not that the platform is inherently malicious. It is that an attacker can steer a legitimate generative system into producing content, code, or workflow output that serves an external objective. In practice, that may involve prompt shaping, indirect instruction, or iterative refinement until the output is fit for phishing, fraud, or malicious scripting. The platform remains trusted infrastructure, but the trust is redirected. That is why simple content filters are insufficient: the abuse happens through normal product capabilities that are difficult to separate from approved use without deeper governance and telemetry.

Practical implication: monitor AI tool access, prompt patterns, and downstream usage for signs of abuse rather than relying only on output moderation.

Why phishing and fraud get harder to distinguish when AI is involved

AI-assisted abuse compresses the difference between content generation and campaign execution. A phishing lure no longer has to be manually written at scale, and a fraud campaign no longer has to be assembled as a one-off operation. The result is higher volume, better variation, and lower operational cost for the attacker. For defenders, that means traditional indicators such as grammar mistakes or repetitive phrasing become less useful. The more important questions are provenance, intent, and whether the account or workflow driving the AI session is behaving consistently with its expected role.

Practical implication: shift detection toward provenance, behavioural anomalies, and abuse of trusted accounts rather than content quality alone.

Identity and trust assumptions behind AI-enabled deception

The governance failure exposed here is that many organisations still assume a legitimate account or approved AI session is synonymous with legitimate intent. That assumption breaks down when the same identity can be used to generate fraudulent content, orchestrate social engineering, or scale malicious automation. In identity terms, access is necessary but not sufficient as a trust signal. Organisations need to treat AI interactions as governed sessions with context, policy, and traceability, especially where human users, service accounts, and AI-assisted workflows intersect.

Practical implication: bind AI access to strong identity context, logging, and policy controls so approved access cannot be treated as implicit trust.


NHI Mgmt Group analysis

Adversarial AI turns trusted workflows into abuse infrastructure: The core issue is not content generation alone, but the repurposing of legitimate AI sessions for fraudulent objectives. Once an approved workflow can be steered into deception, the control problem moves from content review to trust governance. Practitioners should treat AI-enabled abuse as an identity and workflow integrity issue, not a narrow safety-filter problem.

Provenance matters more than surface quality: AI-assisted phishing and fraud reduce the value of traditional content heuristics because the output can be polished, varied, and context-aware. That means defenders cannot depend on obvious language cues to separate benign from malicious activity. The practical implication is that detection and policy must follow the session, the actor, and the downstream use case.

Identity is the missing control plane for AI abuse: Organisations often assume a legitimate login or approved application context is enough to establish trust. That assumption fails when the same identity can generate phishing lures, scripts, or fraud content at scale. The implication is that identity governance for AI must cover session purpose, traceability, and abuse boundaries, not just authentication.

Human, machine, and AI-assisted workflows now share one fraud surface: The convergence here is that a human attacker can use AI platforms, automated services, and trusted accounts together to industrialise abuse. That makes fraud and phishing programmes depend on cross-domain governance rather than isolated email or IAM controls. Practitioners should design for the abuse chain, not only the individual system.

Adversarial AI requires a trust-boundary concept, not just an AI safety concept: The field needs a clearer notion of where legitimate AI use ends and weaponised workflow begins. Without that boundary, organisations will keep overestimating the trustworthiness of approved tools and underestimating how quickly they can be turned into attack infrastructure. Security teams should define and enforce that boundary at the identity, workflow, and telemetry layers.

What this signals

AI abuse is becoming a governance problem, not just a content problem: Organisations that focus only on output moderation will miss the identity, access, and workflow conditions that make legitimate tools exploitable. The stronger control point is the approved session itself, because that is where intent, provenance, and downstream abuse can be constrained.

Human and machine trust signals now need to converge: Fraud teams, IAM teams, and security operations can no longer treat AI-assisted deception as a separate niche. The same control logic that governs privileged access, trusted workflows, and abuse monitoring has to extend into generative AI use where the output can be weaponised quickly.


For practitioners

  • Govern AI sessions as trusted execution paths Classify approved AI tool usage as a governed workflow with identity, purpose, and telemetry attached to each session, rather than as ordinary SaaS activity.
  • Monitor for abuse patterns in prompts and outputs Look for repeated prompt shaping, unusual volume, scripted variation, and downstream content reuse that indicate an AI account is being used for phishing or fraud enablement.
  • Tighten access boundaries around generative tools Limit which users, service accounts, and integrations can call AI systems that can generate text, code, or campaign assets, and log those calls for investigation.
  • Separate productivity use from high-risk workflows Do not allow the same approved workflow to handle ordinary content creation and external-facing deception-prone activity without explicit policy controls and review.
  • Update fraud detection to include AI-assisted content abuse Feed AI-generated campaign indicators into fraud and phishing detection so defenders can respond to higher-volume, higher-variation abuse patterns.

Key takeaways

  • Adversarial AI changes the fraud problem by letting attackers use legitimate platforms as abuse infrastructure for phishing and scripted deception.
  • The practical signal is a shift from content quality checks toward identity, workflow provenance, and suspicious generation behaviour.
  • Security teams should govern AI access as a trusted session problem, because the same approved tools can be repurposed for abuse at scale.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI09 — Human-Agent Trust ExploitationThe article centres on abusing trusted AI workflows for deceptive output.
ASI02 — Tool MisuseAttackers are using legitimate tools for harmful generation and campaign support.
Recommendation — Map AI-enabled deception to ASI09 and restrict trust placed in approved sessions. Apply ASI02 to limit how AI tools can be used in high-risk workflows.
NIST AI RMFGOVERN — AI Governance and AccountabilityThe piece is about governance of legitimate AI systems used abusively.
Recommendation — Establish governance for AI session purpose, ownership, and accountability.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe abuse path depends on who can access and use AI capabilities.
Recommendation — Tighten PR.AA-05 controls around who can access generative AI tools and integrations.
MITRE ATT&CKTA0006;TA0009 — Credential Access; CollectionThe article references phishing and fraud workflows that support adversary operations.
Recommendation — Track AI-assisted deception activity to TA0006 and TA0009 indicators in detections.

Key terms

  • Adversarial AI: Adversarial AI refers to AI used by attackers to scale reconnaissance, impersonation, or abuse in ways that overwhelm normal manual review. For defenders, the issue is not just malicious model use, but the way machine-speed behaviour changes the timing, volume, and accuracy demands on identity and fraud controls.
  • AI-assisted deception: The use of AI systems to create or adapt deceptive content at scale. In practice, it shortens the time needed to craft believable lures, mimic tone, and iterate against targets, which makes social engineering more persistent and harder to recognise across channels.
  • Trust Boundary: A trust boundary is the point where one system’s authority should stop and another system’s authority should begin. For internal automation, weak trust boundaries let monitoring, remediation, and execution share privileges that should have remained separate.
  • Workflow Abuse: Workflow abuse is the use of legitimate business processes such as onboarding, support, or approval chains to gain access that would be harder to obtain through a direct technical exploit. It succeeds when process trust is stronger than identity verification.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity security are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 27, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org