By NHI Mgmt Group Editorial TeamBased on Abnormal AI: “How Abnormal Taught AI Agents To Write Detectors” (May 28, 2026)

TL;DR: AI Detection Agents can turn a customer-reported miss into a deployed detector in hours by analysing attack context, selecting behavioural signals, testing against real traffic, and refining for precision, according to Abnormal AI. The deeper shift is that email defence now depends less on manual review queues and more on whether detectors generalise to attacker intent rather than surface features.


At a glance

What this is: Abnormal AI describes AI Detection Agents that move email defence from manual analyst review to runtime detector creation, using context, behavioural signals, and evaluation loops to catch attack variants.

Why it matters: This matters because IAM, SOC, and email security teams need detections that survive campaign rotation, trusted-platform abuse, and false-positive pressure without relying on brittle keyword rules.


Context

Email defence fails when detection rules are built around surface features that attackers can change cheaply. In this model, the core governance problem is not whether a message looks suspicious, but whether the detector understands the behaviour that makes it malicious in a specific environment.

The article is about AI systems that generate and deploy email detectors from reported misses, which places the problem squarely in the intersection of automation, identity-adjacent trust, and operational security. For practitioners, the issue is whether detection can keep pace with attacker adaptation without turning every miss into a manual review backlog.


Key questions

Q: How should security teams build email detectors that survive campaign variation?

A: They should base detections on attacker behaviour, not on fixed sender or subject values. The rule should capture what makes the message malicious in context, then be tested against historical traffic to confirm it still works when the campaign changes shape. That is how teams avoid brittle rules that only catch the first example.

Q: Why do legitimate-platform phishing campaigns bypass traditional controls so often?

A: They bypass traditional controls because many defences assume malicious delivery comes from suspicious infrastructure. When attackers use trusted SaaS accounts, compromised brands, or approved sending services, reputation-based filters become less reliable. Organisations need account behaviour monitoring, identity verification, and post-delivery detection to catch what gateway controls cannot see.

Q: What are the signs that an automatically generated detector is too narrow?

A: It matches on exact sender domains, reworded subjects, or other surface details that attackers can rotate cheaply. If a detector misses the same campaign once those visible features change, it is memorising examples rather than recognising the underlying attack behaviour. That is a precision problem and a resilience problem at the same time.

Q: How should analysts govern AI-assisted detection deployment?

A: Analysts should require a documented explanation of the behavioural logic, plus evaluation against real traffic before rollout. The control is not just generation speed. It is whether the detector can prove statistical separation, semantic relevance, and low false-positive impact before it reaches production.


Technical breakdown

Behavioural detection versus surface matching

The central technical shift is from pattern matching on fixed indicators to behavioural detection that reasons about attacker intent. Surface features such as sender domain, subject wording, or message formatting are easy to rotate, so they only catch the exact example the model has already seen. Behavioural detection instead asks what is true about the message in context, such as novelty, unusual delivery path, or deviations from normal communication patterns. That makes the detector more resilient to campaign variation and less dependent on static signatures that age quickly.

Practical implication: build detectors around behaviours that survive sender, subject, and delivery-path changes.

Second-order thinking in detector generation

The article’s second-order thinking is the move from raw values to characteristics of those values. A new sender domain is not the useful signal by itself; the useful signal is that the domain is newly registered, unfamiliar, or otherwise atypical for the recipient environment. This abstraction layer is what lets a detector generalise beyond the training example. It also explains why generative models fail when left alone: they often lock onto distinctive surface patterns instead of the underlying condition that made the message suspicious.

Practical implication: validate that each rule maps to a transferable property, not just a one-off indicator.

Why authentication is not enough in trusted-platform abuse

Trusted-platform abuse is the hard case because SPF and DKIM can pass while the message still delivers phishing content through legitimate infrastructure. In that scenario, authentication proves only that the sender used an allowed platform, not that the content or intent is safe. The article’s example shows why behavioural signals matter more than authentication checks alone when attackers hide inside Microsoft Teams event flows or similar approved channels. This is a control boundary problem, not a transport problem.

Practical implication: treat successful authentication as a narrow trust signal, not as proof that the message is benign.


Threat narrative

Attacker objective: The attacker wants to deliver phishing or impersonation content that survives authentication checks and evades static email rules long enough to reach the target.

  1. Entry occurs when an attacker uses trusted-platform abuse, such as a Microsoft Teams event registration flow, to deliver phishing content through legitimate infrastructure.
  2. Credential or trust bypass occurs because SPF and DKIM can pass even though the message content is malicious, allowing the campaign to inherit platform legitimacy.
  3. Impact follows when the next campaign variant changes sender details or subject wording, bypassing brittle rules that matched only the original example.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Behavioural abstraction is now the decisive email security control: detectors that learn the attacker’s method rather than the campaign’s visible shape can survive domain rotation, subject rewriting, and delivery-path changes. That is the real shift from review to runtime. For practitioners, the governance question is whether detection logic is designed to recognise intent or merely replay known examples.

Statistical separation and semantic relevance must both be present: a signal that is statistically different but semantically meaningless, such as a missing attachment in a threat class where attachments are already rare, creates noise rather than protection. Good detection requires the attribute to distinguish malicious from normal traffic and to make sense for the attack type itself. Practitioners should judge automated detection on both dimensions, not on volume alone.

Trusted-platform abuse exposes a trust boundary problem, not an email formatting problem: when SPF and DKIM pass inside legitimate infrastructure, the real failure is over-trusting the transport layer. The message can still be malicious because the platform is being used as a delivery vehicle. The implication is that email defence has to evaluate behaviour across identity, content, and context, not just mail authentication results.

Runtime detector generation changes the operating model for security teams: the most useful output is no longer a queue of analyst-reviewed misses, but a governed pipeline that analyses, tests, refines, and deploys protections continuously. That compresses response time, but it also raises the bar for explainability and evaluation discipline. Practitioners should treat detector generation as an operational control plane, not as a convenience feature.

Identity-adjacent security is moving toward intent-based enforcement: the same logic that governs non-human identities applies here in a different form. If an identity or platform can be legitimately used while still carrying malicious intent, then the control must interrogate behaviour, not just authenticity. Email security teams should expect this model to spread into adjacent trust systems.

From our research library:

What this signals

Behavioural email defence: the useful unit of analysis is no longer the suspicious message alone, but the attacker method that the message expresses. When detection rules are built from that layer, campaign rotation becomes far less effective and the analyst queue becomes a validation point rather than the primary defence line.

Evaluation now governs trust: automated detector generation only makes sense if each rule is stress-tested against real traffic before deployment. The operational implication is that precision, explainability, and rollback discipline matter as much as detection speed.


For practitioners

  • Prioritise behavioural signals over surface features Weight novelty, abnormal delivery context, and communication deviation above sender strings or subject keywords so detectors survive attacker rotation.
  • Validate statistical and semantic fit together Reject detector attributes that separate examples numerically but do not actually describe the attack type in the recipient environment.
  • Test generated detections against broad real traffic Use historical traffic evaluation to surface false positives that only appear outside the initial attack sample, then tighten boundaries before deployment.
  • Treat SPF and DKIM as narrow trust inputs Do not let successful authentication override contextual review when legitimate platforms are being used to carry malicious content.
  • Document why each detector generalises Capture the behavioural rationale for deployment so analysts can trace how the rule maps back to the attacker method, not just the original message.

Key takeaways

  • Email defenders lose ground when rules are tied to sender strings, subject lines, or other surface features that attackers can change quickly.
  • The article’s core method is to reason from behaviour and context, then verify detectors against real traffic before deployment.
  • Security teams should treat successful authentication and fast generation as inputs, not proof that a detection rule is ready for production.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST AI RMF and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI02 — Tool MisuseThe article centers on AI agents selecting signals and deploying detections.
Recommendation — Constrain agent actions to approved detection workflows and verify outputs before deployment.
NIST AI RMFGOVERN — AI Governance and AccountabilityThe article describes governed analysis, testing, and deployment of AI-generated detections.
Recommendation — Define approval, testing, and accountability requirements for AI-driven security outputs.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe article depends on controlling what detection logic can do once it is generated.
Recommendation — Limit automated detection systems to least-privilege actions and review their effective scope.
MITRE ATT&CKTA0005;TA0006 — Defense Impairment; Credential AccessThe campaign pattern abuses trusted platforms to evade detection and deliver phishing.
Recommendation — Map trusted-platform abuse campaigns to ATT&CK and prioritize detections for credential-driven phishing paths.

Key terms

  • Behavioral Detection: A monitoring approach that looks for unusual activity rather than relying only on static inventories. For SaaS integrations, it detects drift in token use, data movement, timing, and endpoint behavior so teams can spot compromise, misuse, or automation that no longer matches its expected pattern.
  • Second-order thinking: A method of analysis that asks what makes a signal suspicious, not just whether the signal is present. It shifts the focus from raw values to the underlying property that continues to hold across variations, which is essential when adversaries can change surface details quickly.
  • Trusted-Platform Abuse: The use of legitimate collaboration or cloud-sharing services as part of the attack infrastructure. The platform itself may be normal business software, but attackers exploit the trust users and security tools place in it to deliver content, redirect victims, or hide malicious activity.
  • Detector Generalisation: Detector generalisation is the ability of a rule or model to identify future variants of the same attack rather than memorising the original example. A generalising detector captures the underlying attacker method, which is essential when campaigns change sender, wording, or delivery path.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 27, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org