By NHI Mgmt Group Editorial TeamBased on Abnormal AI: “EvilTokens: Turning OAuth Device Codes into Full-Scale BEC Operations” (April 3, 2026)

TL;DR: EvilTokens productizes Microsoft 365 compromise by abusing Device Code OAuth, replaying refresh tokens for up to 90 days, and using LLaMA to turn inboxes into BEC intelligence, according to Abnormal AI. The real lesson is that MFA success does not equal session safety when token replay and conditional access gaps remain open.


At a glance

What this is: This analysis shows how EvilTokens uses device code phishing to bypass MFA, steal OAuth refresh tokens, and automate business email compromise against Microsoft 365 accounts.

Why it matters: It matters because IAM teams often treat MFA success as the control outcome, but token replay, conditional access gaps, and session persistence can leave access live after authentication appears complete.


Context

Device code phishing is a credentialless phishing pattern that uses a legitimate OAuth flow instead of a fake login page. The target authenticates on the real identity provider page, which means MFA may complete successfully even though the resulting session is attacker-controlled.

For IAM and NHI programmes, the issue is not password theft but delegated token abuse and session persistence. Once a refresh token is captured, the attacker can keep reusing the granted access until revocation, policy enforcement, or token expiry breaks the chain.

Abnormal AI’s analysis describes a productised phishing-as-a-service operation rather than an isolated campaign. That makes the governance problem broader than user awareness: it is about how identity controls behave after authentication, not just at the point of sign-in.


Key questions

Q: What breaks when device-code phishing is allowed in a Microsoft tenant?

A: The control break is that a legitimate user login can still produce attacker-owned tokens. If device-code flow is broadly available, the phishing path bypasses password theft entirely and turns the approval step into the compromise point. Security teams should treat unrestricted device-code use as a governance exception, not a normal sign-in pattern.

Q: Why do refresh tokens create persistent access after MFA has succeeded?

A: Refresh tokens are designed to mint new access tokens without prompting the user again. When attackers steal them, they can keep reauthenticating until revocation or expiry stops the chain. That is why MFA alone does not close the risk. The dangerous asset becomes the token, not the password.

Q: How can security teams tell whether token replay controls are actually working?

A: They should look for fast revocation of suspicious sessions, short-lived token reuse windows, and conditional access policies that block repeated exchanges from untrusted infrastructure. If tokens remain usable long after the first sign of compromise, the control is failing even when sign-in logs show successful MFA.

Q: What should teams do when a compromised email account is detected?

A: Teams should move immediately from detection to containment. That usually means logging the user out, terminating active sessions, forcing password resets, and adding the user to watchlists that can trigger endpoint containment or reauthentication. The response should be coordinated across identity, email, and endpoint controls so the attacker loses both access and persistence.


Technical breakdown

Device code phishing turns legitimate sign-in into the attack surface

Microsoft’s Device Code OAuth flow is designed for devices that cannot easily host a browser. The user visits a legitimate Microsoft login page, enters a short code, and completes authentication on the provider’s own domain. In EvilTokens, that legitimate flow is abused so the victim completes real MFA while the attacker retains the code exchange context. The phishing page does not need to collect a password or host a fake identity provider. It only needs to steer the user into authorising the attacker-controlled device context. That makes traditional anti-phishing controls less effective because the credential event itself is genuine.

Practical implication: treat device code OAuth as a high-risk sign-in path and restrict it unless headless-device use is explicitly required.

Refresh token replay extends access long after the original login

OAuth refresh tokens are designed to mint new access tokens without repeated user prompts. That is useful for usability, but it also means the original sign-in event can outlive the immediate session. In this case, captured tokens are replayed through infrastructure that automates token exchange twice daily, which keeps mailbox access alive for extended periods. The control failure is not MFA itself but the assumption that successful MFA ends the security problem. Once refresh tokens are stolen, the attacker no longer needs to phish again to maintain access, and conditional access enforcement becomes the key line of defence.

Practical implication: monitor and revoke token grants, not just passwords, and verify that conditional access policies constrain token reuse.

AI turns compromised mailboxes into BEC intelligence in seconds

The post-compromise phase is where the operation becomes more efficient. After mailbox access is established, AI summarisation extracts payment data, account numbers, routing numbers, and wire instructions from inbox contents. That compresses the reconnaissance stage of business email compromise from hours of manual review to seconds of automated extraction. The important technical point is that AI is not being used to create the lure. It is being used to process the victim’s own communications into structured fraud intelligence. That shifts the attacker’s bottleneck from reading email to operationalising it, which is a material change in BEC throughput.

Practical implication: assume mailbox content itself can be machine-processed for fraud intelligence once access is gained, and prioritise rapid session containment.


Threat narrative

Attacker objective: Steal durable Microsoft 365 access and convert mailbox content into fraud-ready intelligence for business email compromise.

  1. Entry occurs through a legitimate Microsoft Device Code OAuth flow, where the victim is guided to authenticate on the provider’s real login page.
  2. Credential authority is converted into attacker-controlled access when OAuth refresh tokens are captured and exchanged for reusable mailbox sessions.
  3. Persistence is maintained by replaying those tokens twice daily, keeping the compromised Microsoft 365 account usable for up to 90 days.
  4. Impact is realised when inbox contents are mined for payment details and wire instructions that support business email compromise.
  • Microsoft Midnight Blizzard breach: Midnight Blizzard (APT29) exploited legacy test account without MFA to breach Microsoft.
  • Uber breach 2022: A contractor's stolen password and MFA fatigue gave a Lapsus$-linked attacker Uber's internal tools; Uber rotated keys to many services.

Read and download The State of NHI & AI Agent Breach Report 2026, covering 150+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

MFA completion is not an access-ending event: This case shows that successful authentication can still produce attacker-controlled sessions when the resulting token chain is the real target. The governance mistake is treating MFA as the control boundary instead of one step in a larger authorisation and session-risk model. IAM teams need to evaluate whether their sign-in controls actually govern the post-authentication state, not just the login event.

Device code phishing exposes a session-governance gap, not a user-training gap: The victim authenticates on the real Microsoft page, so the failure is not a fake-login detection problem. The broken premise is that only suspicious credentials create risk, when in fact legitimate OAuth flows can be redirected into attacker use. That means identity governance has to examine which flows should exist at all, especially when headless-device use is not a business requirement.

Refresh token replay creates token durability that outlives user intent: Once tokens are issued, the attacker’s access can persist independently of the user’s awareness or the original prompt. That is an NHI-style governance problem because the credential behaves as a durable bearer artefact rather than a one-time proof of interaction. The practical implication is that lifecycle control over tokens matters as much as the original authentication ceremony.

AI-assisted inbox exploitation reduces the attacker’s operational cost of BEC: The meaningful change is not that a model is present, but that mailbox review becomes machine-assisted extraction. That compresses the time between compromise and fraud execution and makes access more valuable at lower scale. Practitioner teams should assume that any successfully accessed mailbox may now be rapidly mined for structured financial data, not just read by a human operator.

Device code OAuth needs category-level governance, not edge-case exception handling: This attack pattern shows how a rarely reviewed authentication path can become a high-yield abuse channel when it is left enabled by default. The issue is not whether device code works for legitimate headless devices, but whether organisations have formally bounded where it is allowed and why. The implication is a stricter inventory of non-interactive sign-in paths across the identity estate.

From our research library:

What this signals

Device code OAuth has become a category-level abuse path: Organisations should not treat this as a one-off phishing trick. When a legitimate identity flow can be used to complete MFA and hand over token authority, the question becomes whether the flow belongs in your environment at all.

Token lifecycle control now sits beside MFA in the control stack: The attack only becomes durable when refresh tokens remain valid long enough for replay. That pushes teams to govern issuance, revocation, and session evaluation as a single identity-control problem rather than separate tasks.

Mailbox compromise is increasingly an automation problem, not just an access problem: Once an attacker can read mailboxes at scale, AI-assisted extraction can turn a small number of compromised accounts into high-confidence fraud targets. The programme implication is that finance and executive mailbox protection needs to be treated as an identity risk as well as a data risk.


For practitioners

  • Disable device code authentication where it is not required Remove this OAuth flow through Conditional Access for environments that do not need headless-device sign-in, so the attacker loses the code-based entry path entirely.
  • Treat refresh tokens as governed credentials Review token issuance, reuse, and revocation behaviour alongside passwords and MFA, because captured refresh tokens can preserve mailbox access after the initial sign-in.
  • Shorten the time between token compromise and revocation Enable continuous access evaluation and confirm that policy enforcement reaches active sessions quickly enough to interrupt replay before persistence is established.
  • Limit mailbox exposure to fraud-enabling data Apply tighter controls around executive and finance mailboxes, since AI-assisted review of inbox content turns ordinary correspondence into wire-fraud intelligence.

Key takeaways

  • Device code phishing can bypass the assumptions many teams make about MFA because the victim authenticates on a real Microsoft page, not a fake one.
  • Refresh token replay turns a single successful sign-in into durable access that can survive far beyond the original authentication event.
  • The practical control response is to restrict device code OAuth, govern token lifecycle more tightly, and shorten the window in which replayed tokens remain useful.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST SP 800-53 Rev 5 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-04 — Insecure AuthenticationDevice code phishing abuses a legitimate authentication flow to obtain tokens.
NHI-07 — Long-Lived SecretsRefresh tokens remained usable long enough to preserve mailbox access for weeks.
NHI-02 — Secret LeakageCaptured OAuth tokens are the secret material that powers replay after initial compromise.
Recommendation — Restrict device code authentication paths and review where non-interactive sign-in is genuinely required. Shorten token lifetimes and revoke durable tokens as soon as suspicious reuse is detected. Treat token theft as secret leakage and rotate or revoke exposed credentials immediately.
NIST SP 800-53 Rev 5IA-5 — Authenticator ManagementAuthenticator lifecycle governs issuance, revocation, and reuse of tokens in this attack.
Recommendation — Apply authenticator management to control token issuance, reuse, and revocation across the tenant.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe attack succeeds when granted access remains usable beyond the intended session boundary.
Recommendation — Review authorisations that allow repeated token exchange and remove unnecessary standing access paths.
MITRE ATT&CKTA0006;TA0010 — Credential Access; ExfiltrationThe campaign steals tokens and uses mailbox access to collect fraud-enabling information.
Recommendation — Map token theft and mailbox harvesting to credential access and exfiltration detections in your monitoring.

Key terms

  • Device Code OAuth Flow: An OAuth sign-in method for devices that cannot easily host a browser or keyboard. A user enters a short code on a legitimate identity provider page, which makes it attractive for attackers because the authentication event can be real even when the resulting session is abused.
  • Refresh Token: A longer-lived credential that can mint new access tokens without forcing the user to authenticate again. Because refresh tokens can preserve access for extended periods, they are a major governance concern when malicious or over-scoped applications are granted consent.
  • Token Replay: Token replay is the reuse of a valid access or refresh token by someone other than the intended client. The token may still be unexpired and cryptographically correct, so the compromise often shows up only through context anomalies such as location, device, or session overlap.
  • Business email compromise: A form of social engineering where an attacker impersonates a trusted person or domain to manipulate payment, change banking details, or extract sensitive information. It often succeeds without malware because the attacker targets process trust and human judgement instead of technical controls.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 27, 2026.
Updated on October 8, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org