By NHI Mgmt Group Editorial TeamBased on Abnormal AI: “Threats without Borders: Insights from the Stockholm AI Roadshow” (June 8, 2026)

TL;DR: Anticimex blocked more than 40,000 malicious emails that Microsoft missed between February and April, avoiding an estimated $169,000 in losses through AI Security Mailbox automation and graymail filtering, according to Abnormal AI. The bigger lesson is that email defence built for older threat volumes and patterns is no longer keeping pace with AI-accelerated attacks.


At a glance

What this is: This is an analysis of how AI-powered email attacks are slipping past native Microsoft protections, with Anticimex cited as a case where more than 40,000 malicious messages were missed before Abnormal was deployed.

Why it matters: It matters because email remains a primary identity attack path, and IAM, PAM, and security teams need controls that can detect behaviour and scale beyond static, rule-based filtering.


Context

AI-powered email attacks are messages that use automation and language generation to become more convincing, more personalised, and easier to scale than older phishing campaigns. In this article, the governance problem is not just email security, but the mismatch between modern attack volume and native Microsoft protections that were built for a different threat profile.

Abnormal AI uses Anticimex as the example: employees spotted near-misses that the native controls did not catch, then the company deployed additional behavioural filtering and mailbox automation. The broader issue for IAM and security leaders is that email compromise often becomes the first step in account takeover, fraud, or privilege abuse.


Key questions

Q: Why do native email protections miss AI-powered phishing campaigns?

A: Native email protections often depend on known indicators, reputation, and repeatable patterns. AI-powered phishing changes wording, timing, and context faster than those controls can adapt, so messages can look legitimate enough to pass. The failure is usually a speed and scale mismatch, not a complete absence of filtering.

Q: How should security teams respond when AI makes business email compromise harder to spot?

A: Teams should move beyond message inspection and verify the requester, the channel, and the business context before allowing action. AI makes tone and wording unreliable signals, so the control point becomes workflow validation, out-of-band confirmation, and monitoring for abnormal approval patterns across finance, executive, and supplier interactions.

Q: What signs show that email identity controls are not keeping pace?

A: Watch for stale shared mailboxes, persistent delegated send permissions, accounts that remain active after role change, and alerts that show unusual sending behaviour from trusted identities. Those signals usually mean the organisation is protecting the inbox but not the identity behind it.

Q: What should teams do after malicious email reaches users despite native protection?

A: Teams should review the message path, identify which detection stage failed, and connect mailbox events to account and identity monitoring. The point is to contain the exposure chain, not just delete the email after the fact. If the message was credible enough to trigger interaction, response should extend beyond the inbox.


Technical breakdown

Why native Microsoft email controls miss modern attacks

Native email protections typically rely on known indicators, reputation signals, and policy logic that work well against repeatable threats. AI-generated phishing and business email compromise change the content faster than static rules can adapt, and they often blend into normal business language. That creates a detection gap where the message looks ordinary enough to pass basic filtering but still carries malicious intent. In practice, the control failure is not a lack of email security technology, but a mismatch between attacker adaptation speed and defender model update speed.

Practical implication: tune detection for behavioural anomalies and escalation patterns, not just signature-style indicators.

AI Security Mailbox automation and graymail filtering

Mailbox automation in this context means applying policy-driven triage to message flows so suspicious or low-value mail is handled before a user has to make the decision. Graymail filtering reduces noise by separating unwanted but non-malicious mail from truly risky content, which matters because alert fatigue can hide the one message that matters. The article’s point is that both functions improve resilience by changing what reaches the human inbox, not by assuming users will spot every social engineering attempt.

Practical implication: reduce inbox exposure and review burden so users are not the last line of defence against AI-crafted email.

Behavioral AI as a response to scaled social engineering

Behavioral AI looks for patterns in sender behaviour, conversation context, and message intent rather than relying only on whether an attachment or URL is already known to be bad. That matters when attackers generate many slightly different messages at machine speed, because each individual email may look unique while the campaign pattern remains consistent. For identity teams, this matters because email is often the first credential-adjacent control point, and failure there can cascade into account compromise, token theft, or downstream privilege abuse.

Practical implication: correlate mail signals with identity risk signals so email findings can drive faster access review or containment.


Threat narrative

Attacker objective: The attacker aims to get malicious email into the inbox, exploit user trust, and create a path to account compromise or fraud at scale.

  1. Entry occurs through AI-powered email messages that mimic legitimate business communication closely enough to bypass native filtering and reach users.
  2. Credential or trust abuse follows when users engage with messages that appear authentic, opening the door to compromise of accounts or workflows.
  3. Impact comes from missed malicious messages accumulating at scale, increasing the probability of fraud, account takeover, or other downstream identity compromise.
  • DeepSeek database exposure 2025: An unauthenticated DeepSeek ClickHouse database exposed over a million log lines with plaintext chat history and API keys in 2025.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Email security built for static threat volumes is now a governance problem, not just a filtering problem. AI-generated attacks shorten the defender’s decision window and increase the number of messages that must be evaluated. That shifts the burden from user judgement to control design, especially where email is an upstream identity control point. Security leaders should treat inbox filtering as part of identity risk management, not an isolated mail issue.

Behavioral detection is becoming the dividing line between usable inboxes and exploitable inboxes. The article shows that native Microsoft protections and vigilant employees still left exposure, which means the issue is not awareness alone. The practical standard is whether a control can separate legitimate business communication from machine-scaled social engineering at production volume. Teams that cannot do that will keep relying on users to spot what systems missed.

AI-powered email attacks create identity blast radius before any password is touched. A convincing message can trigger trust, workflow changes, or credential capture without needing direct malware delivery. That means email defence, identity telemetry, and privilege controls have to be treated as a linked chain. Practitioners should assume the first compromised resource may be the human trust layer, not the endpoint.

Graymail reduction is now a security control, not an inbox convenience feature. When low-value mail floods the same channel as malicious mail, the real risk is signal dilution. Anticimex’s result suggests that reducing noise can materially improve the chances of spotting genuine attacks. Security programmes should treat message volume management as part of control effectiveness, because overload is itself a vulnerability.

Named concept: inbox decision debt. Every extra message a user has to judge manually increases the chance that a sophisticated attack blends into routine communication. AI accelerates that debt by making campaigns more personalised and more frequent. The implication is straightforward: the control plane has to move earlier than the human decision point, or the inbox becomes an attacker-managed queue.

What this signals

Inbox decision debt: AI-powered phishing increases the number of messages humans must evaluate, which means the inbox itself becomes a control bottleneck. Teams should assume that user judgement will fail more often as campaigns become more personalised and frequent.

Email security and identity security are converging because one convincing message can now initiate account compromise without malware. Practitioners should watch for cases where mail controls, access monitoring, and response processes remain disconnected, because that is where attacker momentum builds.


For practitioners

  • Strengthen behavioural email detection Prioritise message analysis that looks at sender behaviour, conversation patterns, and abnormal intent rather than only known bad indicators.
  • Reduce inbox noise with graymail controls Separate low-value mail from genuinely risky mail so analysts and users are not forced to inspect every message at full volume.
  • Treat email as an identity risk signal Feed suspicious mail events into account monitoring, access review, and response workflows so phishing and compromise signals are not handled in isolation.
  • Reassess native Microsoft protection boundaries Measure where native controls let malicious mail through and define the conditions under which layered mail defence is required.

Key takeaways

  • AI-powered email attacks are now outrunning controls built for static threat patterns and lower message volume.
  • The Anticimex example shows that native Microsoft protections can miss large numbers of malicious emails at operational scale.
  • Security teams need earlier triage, behavioural detection, and tighter linkage between email events and identity response.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 sets the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10ASI09 — Human-Agent Trust ExploitationAI-generated email attacks exploit human trust at machine speed.
Recommendation — Map mailborne social engineering to trust exploitation patterns and tighten pre-click verification controls.
MITRE ATT&CKTA0001;TA0006 — Initial Access; Credential AccessThe article centers on phishing as the entry path to identity compromise.
Recommendation — Hunt email-delivered initial access and credential capture together rather than treating them as separate problems.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsMailborne attacks become identity risk when they drive unauthorised access attempts.
Recommendation — Link suspicious-email handling to identity authorization review and containment workflows.

Key terms

  • Behavioural email detection: A detection approach that looks for patterns in sender behaviour, message timing, language change, and downstream user interaction rather than relying only on signatures. It is designed to catch attacks that mutate quickly. For identity programmes, its value is in finding the moment an email becomes an access risk.
  • Graymail: Graymail is legitimate but low-value email that competes with important messages for attention. In security operations, it matters because it lowers signal quality, makes anomalous mail easier to miss, and can degrade the effectiveness of both human review and behavioral detection.
  • Inbox Decision Debt: The growing risk created when users must personally judge too many messages for legitimacy. The more AI-driven campaigns increase frequency and personalization, the more that decision burden compounds, making human review less reliable as a primary control.
  • Mailbox Automation: Automated handling of suspicious, repetitive, or low-value email events, including quarantine, cleanup, and filtering actions. Used correctly, it shortens exposure windows and reduces dependence on end users as the final control in a phishing chain.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 27, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org