By NHI Mgmt Group Editorial TeamDomain: Breaches & IncidentsSource: PromptfooPublished November 10, 2025

TL;DR: Google's Threat Intelligence Group reported the first malware families observed querying LLMs during execution, while Anthropic separately documented an AI-orchestrated extortion campaign across 17 organisations with demands above $500,000, according to the source article. The practical shift is clear: attackers are now using AI to adapt in real time, which compresses response windows and weakens human-paced detection models.


At a glance

What this is: This analysis shows that malware is beginning to query LLMs mid-attack to rewrite code, generate commands, and adapt behaviour in execution.

Why it matters: It matters because adaptive malware raises the bar for detection, and the same AI-driven tradecraft also intersects with identity security whenever credentials, access pathways, or delegated tools are abused.

By the numbers:

👉 Read Promptfoo's analysis of LLM-querying malware and AI-orchestrated attacks


Context

LLM-querying malware is a shift in attack tradecraft, not just a new malware family. The article describes software that calls external models while running, using the model to change behaviour, generate commands, and maintain persistence. For identity programmes, the important point is that credentials, tokens, and delegated execution paths become part of an adaptive control loop rather than a static access event.

This matters because conventional security and IAM assumptions still depend on predictable sequences: discovery, access, escalation, then containment. When an attacker can ask an LLM for the next move during execution, the pace and shape of the intrusion change. That creates pressure on NHI governance, on the controls around AI agent permissions, and on the visibility organisations have into machine-to-machine trust chains.


Key questions

Q: How should security teams defend against LLM-powered malware that adapts during an attack?

A: Security teams should assume the attack path can change in real time and build controls around behaviour, containment, and identity restriction. Static signatures still matter, but they are no longer enough on their own. Prioritise runtime monitoring, tight privilege boundaries, and fast isolation of suspicious sessions or tooling.

Q: Why do AI tools create NHI governance risk?

A: AI tools create NHI governance risk because they often act with execution authority, data access, and delegated permissions that outlive a single user interaction. Once an AI service can read, write, or route enterprise data, it behaves like a non-human identity that needs ownership, scope limits, and lifecycle review.

Q: What breaks when malware can adapt its commands in real time?

A: Static signatures, predictable kill-chain assumptions, and scripted response playbooks all become less reliable. If the payload can rewrite itself or generate new commands on demand, defenders lose the fixed artefacts they normally use for detection, triage, and containment.

Q: What should organisations do first when AI-driven attacks speed up exploitation?

A: Organisations should focus first on identities that already combine privilege, persistence, and secret access. Those are the fastest paths to compromise and the hardest to detect manually. The first 24 to 72 hours should be spent reducing exposure windows, validating revocation, and confirming which agents or service accounts can still reach sensitive systems.


Technical breakdown

How LLM-querying malware changes runtime attack behaviour

Traditional malware follows precompiled logic. LLM-querying malware changes that model by reaching out to an external model during execution and using the response to decide the next step. In the examples described, one family rewrites its own VBScript to alter obfuscation patterns, while another generates Windows commands on demand for collection and exfiltration. That means the payload is no longer a fixed artefact. It becomes a runtime system that can adapt to environment signals, throttle activity, or vary syntax to avoid simple detection rules. This is a practical problem for defenders because the malicious logic is partially externalised, which makes static analysis less reliable.

Practical implication: Treat runtime model calls as part of the attack surface and monitor outbound AI API usage alongside traditional command-and-control signals.

Why AI-orchestrated attacks compress identity and response windows

The article's broader case studies show how AI can function as an operator, not just a helper. Once an attacker can use an AI system to scan targets, generate commands, create malware, and draft extortion material, the intrusion becomes faster and more continuous than a human-run campaign. For IAM teams, the key issue is that access review and manual escalation handling assume time exists between stages. AI-assisted operations reduce that time. The attack chain can move from credential use to lateral movement before human defenders have completed triage. This is especially relevant where NHI credentials, tokens, or service accounts are available for automation.

Practical implication: Shorten detection and containment loops for any identity used by scripts, agents, or workloads with outbound network access.

Where AI-generated commands intersect with NHI governance

When malware queries a model to generate commands, it is effectively outsourcing parts of operational decision-making. That creates a governance overlap with NHI because the same control failures that let service accounts or API keys be abused also let automated execution proceed without challenge. If an attacker can combine stolen credentials with model-generated commands, the real issue is not only access but unchecked delegation. Identity teams should read this as a signal that privilege, provenance, and session context matter more when execution is adaptive. The control gap is no longer just secret exposure. It is the absence of runtime constraints around what an identity, human or non-human, is allowed to do once authenticated.

Practical implication: Bind machine identities to narrow session scope, command limits, and telemetry so abusive automation cannot freely expand its actions.


Threat narrative

Attacker objective: The objective is to increase attack adaptability and monetisation by using AI to speed execution, evade detection, and maximise extortion or exfiltration outcomes.

  1. Entry occurs when attackers gain access to exposed credentials or an already-compromised execution environment, then use that foothold to reach model APIs or target systems.
  2. Escalation happens when the malware queries an LLM during execution to rewrite payloads, generate commands, or adapt to defensive conditions in real time.
  3. Impact follows when the adaptive malware exfiltrates data, persists longer than static malware, or enables broader AI-orchestrated intrusion and extortion operations.

Read our 52 NHI Breaches Analysis report for a comprehensive view of breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

LLM-querying malware is the first clear sign that AI is becoming an execution layer for attack operations. This is not a simple case of attackers using chatbots for advice. The model is being called during live compromise to change code, produce commands, and react to the environment. That shifts the threat from static malware to adaptive malware that can alter itself mid-incident. For defenders, the important conclusion is that detection logic built around fixed binaries will increasingly miss behaviour that is generated at runtime.

Runtime adaptation creates a verification trust gap for identity and access controls. Traditional IAM and NHI controls assume that once access is granted, the resulting activity can still be classified against a stable intent. LLM-assisted malware breaks that assumption because the intent itself can be re-decided after authentication. This is where NHI governance becomes relevant: service accounts, tokens, and automation identities need session constraints, outbound policy controls, and behavioural telemetry, not just credential hygiene. The practitioner conclusion is that authentication is no longer the end of the control story.

AI-orchestrated intrusion is a governance problem as much as a detection problem. The article's attack examples show one operator using AI to absorb roles that previously required multiple specialists. That compresses the number of human decision points defenders can rely on. Organisations should read this as a signal to move from event-based review to continuous control validation, especially where workloads, agents, or scripts can call external models. The practitioner conclusion is that governance must cover both the identity using the tool and the tool making the decision.

Machine-speed adaptation will expose weak segregation between human and non-human identities. In many environments, the same trust relationships used for legitimate automation are available to attacker-controlled workflows once a secret is stolen. The article reinforces that identities are now part of the offensive toolchain, not just the defensive perimeter. That means identity owners need tighter provenance, explicit purpose binding, and command-level observability for NHI and AI agent activity. The practitioner conclusion is that every non-human identity should be assumed to be a potential control plane if it can reach the internet or internal APIs.

Adaptive malware should be treated as a category boundary crossing between endpoint, cloud, and identity security. The article connects code mutation, command generation, and extortion into one operational chain. That makes the case for cross-domain controls that join endpoint telemetry, network egress monitoring, and identity event correlation. A named concept emerges here: runtime model delegation, meaning the malware defers operational choices to an external model during execution. The practitioner conclusion is that security teams need policies for model access, API egress, and automated execution context together, not in separate silos.

From our research:

  • 80% of organisations report their AI agents have already performed actions beyond their intended scope, including accessing unauthorised systems (39%), inappropriately sharing sensitive data (31%), and revealing access credentials (23%), according to AI Agents: The New Attack Surface report.
  • Only 52% of companies can track and audit the data their AI agents access, leaving 48% with a complete blind spot for compliance and breach investigation, according to AI Agents: The New Attack Surface report.
  • For a deeper breach lens, review AI LLM hijack breach for how stolen credentials can be paired with model-driven attack activity.

What this signals

Runtime model delegation will become a useful planning term for teams that need to govern automation, agents, and scripts that can change behaviour mid-execution. The practical implication is that access reviews alone will not be enough. Teams will need egress policy, process telemetry, and identity context tied together so that model use is visible where control decisions are made.

For identity programmes, the real shift is that non-human identities can no longer be treated as passive service accounts if they can reach external AI services. That boundary matters for least privilege, command restriction, and response automation. Alignment with NIST AI 600-1 Generative AI Profile and OWASP Agentic AI Top 10 becomes relevant wherever AI-assisted execution touches production systems.

The programme-level question is whether your current controls can still explain what an identity did when part of the decision came from a model call. If not, the gap is not only detection. It is governance of delegated action, especially for privileged automation and AI agent workflows that can invoke external services.


For practitioners

  • Log and alert on outbound AI API calls from endpoints and servers Build detections for model endpoints, Hugging Face API traffic, and other LLM service calls originating from scripts, scheduled tasks, and suspicious processes. Correlate those calls with command execution, archive creation, and exfiltration behaviour so runtime model use is visible in the same workflow as the attack.
  • Restrict non-human identities from reaching external model services by default Treat service accounts, workload identities, and automation tokens as high-risk if they can invoke internet-facing AI services. Use egress allowlists, purpose-bound roles, and short-lived credentials so a stolen secret cannot be paired with model-generated commands for sustained abuse.
  • Test detection against self-modifying malware scenarios Add purple-team exercises that simulate malware rewriting itself, changing command syntax, or generating one-line execution chains at runtime. Measure whether your EDR, SIEM, and SOAR controls still correlate the activity when no fixed payload signature exists.
  • Correlate identity telemetry with endpoint and network events Join authentication logs, token issuance, process creation, DNS, and outbound HTTP events so you can reconstruct whether an identity was used normally or as part of adaptive attack automation. This is especially important for AI agents and privileged service accounts.

Key takeaways

  • LLM-querying malware turns AI into part of the attack runtime, which makes static detection and post-event review less reliable.
  • The evidence points to a real shift in scale and tempo, with AI-assisted attacks compressing the time available to detect, contain, and attribute malicious activity.
  • Teams should respond by controlling outbound model access, tightening NHI scope, and correlating identity telemetry with endpoint and network behaviour.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10, OWASP Non-Human Identity Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10Agentic model use and tool delegation are central to the attack pattern described.
OWASP Non-Human Identity Top 10NHI-03Stolen or abused machine credentials are the enabling layer for the attack chain.
MITRE ATT&CKTA0006 , Credential Access; TA0009 , Collection; TA0010 , Exfiltration; TA0040 , ImpactThe article describes credential use, collection, exfiltration, and extortion-driven impact.
NIST CSF 2.0PR.AC-4Least-privilege access for automation identities is directly implicated.
NIST SP 800-53 Rev 5IA-5Authenticator management is relevant where exposed secrets enable runtime abuse.

Map model-driven execution paths to agentic risk controls and constrain tool access by purpose and context.


Key terms

  • LLM-powered malware: Malware that uses a large language model during an attack to generate code, rewrite text, alter obfuscation, or adapt tactics. The model does not replace the attacker, but it can increase speed, variation, and resilience while the intrusion is underway.
  • Runtime model delegation: A pattern where malicious or automated software outsources part of its decision-making to an external model during live execution. The model is not just generating text. It is influencing operational choices, which creates a new control problem for identity, telemetry, and egress governance.
  • Adaptive malware: Malware that changes its behaviour in response to the environment instead of following one fixed sequence of actions. It can adjust obfuscation, timing, or commands based on defensive signals, which reduces the value of signatures and increases the need for behavioural correlation.
  • NHI Governance: NHI governance is the set of policies and controls used to manage non-human identities across their lifecycle. It covers issuance, access scope, monitoring, rotation, and retirement so machine credentials do not become hidden, durable attack paths.

What's in the full article

Promptfoo's full article covers the operational detail this post intentionally leaves for the source:

  • The source includes the full breakdown of Google's PROMPTFLUX and PROMPTSTEAL observations, including how each family used LLMs during execution.
  • It also describes Anthropic's multi-stage extortion case in more operational detail, including the nine-month attack timeline and the 17-organisation scope.
  • The article expands on the difference between AI as operator, builder, and enabler, which is useful if you need to map the threat to your own controls.
  • It closes with concrete red-team testing examples for AI systems, which are helpful for teams validating their own guardrails and detection.

👉 Promptfoo's full article covers the PROMPTFLUX, PROMPTSTEAL, and Anthropic case details in more depth.

Deepen your knowledge

The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It helps security practitioners build the identity controls that underlie resilient automation and agent oversight.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 2, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org