Subscribe to the Non-Human & AI Identity Journal

Notifications
Clear all

Cheap AI models for cyber work: what do security teams do now?


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 15374
Topic starter  

TL;DR: Lower-cost models such as Muse Spark 1.1 and GLM are narrowing the gap to premium cyber models, and XBOW’s black-box testing shows they can still rediscover real vulnerabilities when given enough runtime and budget discipline. Good-enough offensive capability is becoming cheap enough to matter, which raises the floor for both attackers and defenders.

NHIMG editorial — based on content published by Xbow: The Rise of Affordable Models: Comparing GLM and Muse Spark on Cyber

Questions worth separating out

Q: What changes when cheaper AI models can support offensive security work?

A: The main change is economic, not just technical.

Q: How should security teams respond to AI models that can iterate cheaply?

A: They should focus on reducing the value and lifespan of exposed access rather than hoping expensive models remain out of reach.

Q: What do organisations get wrong about model benchmarks?

A: Organisations often mistake benchmark scores for trust evidence.

Practitioner guidance

  • Re-score offensive risk by cost per attempt Review AI-related threat models using the number of exploit attempts a cheap model can run before detection, not just the strength of a single model output.
  • Inventory exposed secrets and delegated access paths Map service accounts, API keys, tokens, and other credentials that a model-assisted workflow could probe from the outside.
  • Shorten revocation and rotation cycles Reduce the usable lifetime of credentials that could be abused by repeated AI-driven testing.

What's in the full report

Xbow's full analysis covers the operational detail this post intentionally leaves for the source:

  • Benchmark figures and figure-by-figure comparisons across GLM, Muse Spark 1.1, Mythos, and GPT-5.5 in offensive security tasks
  • The exact evaluation setup used for black-box testing against vulnerable open-source applications
  • Cost and cache-efficiency discussion behind the progress-per-dollar view
  • Model-by-model observations on long-horizon task performance and recovery from dead ends

👉 Read Xbow's analysis of affordable AI models in offensive security →

Cheap AI models for cyber work: what do security teams do now?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 14958
 

Cheap model capability creates an offensive scale problem, not just a quality problem. The article shows that a model does not need to be the strongest available to be operationally relevant. Once it can run long enough, retry often enough, and recover from dead ends, it becomes economically useful for exploitation workflows. For practitioners, that means the threat model shifts from elite capability to affordable repetition.

A question worth separating out:

Q: How can teams know whether AI-assisted probing is becoming a real threat?

A: Look for rising volumes of low-signal exploratory traffic, repeated dead ends, and faster movement from suspicion to reproducible abuse. If your telemetry cannot separate automated search from ordinary usage, the organisation is already underprepared for cheap AI-enabled attack workflows.

👉 Read our full editorial: Affordable AI models are changing the economics of offensive security



   
ReplyQuote
Share: