By NHI Mgmt Group Editorial TeamDomain: AI SecuritySource: XbowPublished July 9, 2026

TL;DR: Lower-cost models such as Muse Spark 1.1 and GLM are narrowing the gap to premium cyber models, and XBOW’s black-box testing shows they can still rediscover real vulnerabilities when given enough runtime and budget discipline. Good-enough offensive capability is becoming cheap enough to matter, which raises the floor for both attackers and defenders.


At a glance

What this is: XBOW’s analysis shows that lower-cost AI models are becoming capable enough to support real offensive security workflows, even if they still trail frontier systems at peak performance.

Why it matters: That matters to IAM and security programmes because the economics of abuse are shifting: more capable, cheaper models can scale vulnerability discovery, automation, and exploit research faster than governance teams can rely on human-review bottlenecks alone.

👉 Read Xbow's analysis of affordable AI models in offensive security


Context

Affordable AI models are lowering the cost of advanced offensive security work, which means the barrier to running useful exploratory testing is no longer reserved for the most expensive frontier systems. In practice, the question for security teams is not whether peak model quality still matters, but whether cheaper models can already create enough scale, persistence, and reach to alter the threat economics for offensive security and identity governance.

XBOW’s testing is a reminder that model capability is only one part of the control problem. What matters for practitioners is how cheaply a model can search, retry, and recover across an attack path. That has an identity dimension wherever the workflow touches credentials, tokens, secrets, or delegated access, because automation lowers the cost of brute-force discovery while increasing pressure on lifecycle controls and runtime authorization.


Key questions

Q: What changes when cheaper AI models can support offensive security work?

A: The main change is economic, not just technical. Lower-cost models make it feasible to run more search, more retries, and more parallel probing for the same budget. That expands the number of actors who can automate discovery and makes defensive assumptions based on attacker cost far less reliable.

Q: How should security teams respond to AI models that can iterate cheaply?

A: They should focus on reducing the value and lifespan of exposed access rather than hoping expensive models remain out of reach. That means tightening secrets scope, revocation speed, and monitoring for repeated exploratory behaviour across internet-facing systems and delegated access paths.

Q: What do organisations get wrong about model benchmarks?

A: Organisations often mistake benchmark scores for trust evidence. Benchmarks can be gamed, tuned, or narrowly optimised, so they do not prove that a model is safe in production or resistant to manipulation. Practitioners should use benchmarks as a screening tool, then require independent verification and reproducibility checks.

Q: How can teams know whether AI-assisted probing is becoming a real threat?

A: Look for rising volumes of low-signal exploratory traffic, repeated dead ends, and faster movement from suspicion to reproducible abuse. If your telemetry cannot separate automated search from ordinary usage, the organisation is already underprepared for cheap AI-enabled attack workflows.


Technical breakdown

How lower-cost models change black-box vulnerability discovery

Black-box offensive testing asks a model to reason from exposed behaviour rather than source code. The workflow is iterative: explore an application, form a hypothesis, test it, and recover when the path fails. Lower-cost models matter because the economics of iteration improve. A model does not need frontier-level reasoning on every step if it can run longer, retry more often, and still reach a reproducible exploit. That shifts the bottleneck from raw intelligence to cost per successful attempt, which is a different security equation altogether.

Practical implication: treat model economics as part of the attack surface, not just model quality.

Why progress per dollar matters more than leaderboard rank

Leaderboard rank measures peak capability under controlled conditions, but real offensive workflows are budgeted and time-bound. A cheaper model that is good enough to keep searching can outperform a stronger model that is too expensive to run at scale. This is especially relevant in cyber operations where the goal is not perfect reasoning, but enough useful progress to move from suspicious behaviour to a working exploit. The article’s core point is that attackers optimise for economics, not prestige metrics.

Practical implication: evaluate AI-enabled security tooling by task completion cost, not benchmark position alone.

Where identity and secrets become the real leverage point

When AI models are used to accelerate offensive work, the highest-value targets are often credentials, secrets, tokens, and delegated access paths. Those are Non-Human Identity assets in operational terms, even when the article is about model performance rather than identity security. Cheaper automation increases the number of attempts against exposed secrets and makes weak lifecycle controls more consequential. The identity issue is not the model itself, but the access patterns it can scale once it is embedded in an offensive workflow.

Practical implication: prioritise secrets governance, access scoping, and revocation speed where AI-assisted abuse can scale.


NHI Mgmt Group analysis

Cheap model capability creates an offensive scale problem, not just a quality problem. The article shows that a model does not need to be the strongest available to be operationally relevant. Once it can run long enough, retry often enough, and recover from dead ends, it becomes economically useful for exploitation workflows. For practitioners, that means the threat model shifts from elite capability to affordable repetition.

Progress-per-dollar is becoming the more useful security metric. Leaderboards tell you which model wins a benchmark, but attackers care about cost, latency, and run length. A cheaper model that can search across many targets may create more real-world risk than a pricier model used sparingly. Security teams should therefore assess AI risk through workload economics, not model prestige.

Identity governance becomes more exposed when offensive automation gets cheaper. The genuine intersection here is with secrets, tokens, and delegated access, because those controls determine what an AI-assisted workflow can touch once it starts probing. Where runtime access is broad and revocation is slow, cheaper models can amplify exposure faster than review processes can contain it. The control gap is not AI novelty, it is weak access lifecycle discipline.

Model capability pressure will spread from frontier systems to broadly available ones. The article’s most important implication is that advanced offensive techniques will not stay confined to premium models for long. That broadens the audience that can attempt semi-automated exploitation and makes control assumptions based on cost scarcity less reliable. Practitioners should assume capability diffusion, then govern for it.

Lower-cost AI will expose maturity gaps in identity and secrets programmes. When attack iteration gets cheaper, every unresolved secret, stale token, and over-broad service account becomes easier to pressure at scale. The organisations that understand their Non-Human Identity estate, and can revoke access quickly, will absorb this shift more safely than those still relying on manual review cycles.

What this signals

Lower-cost models will make automated probing more affordable across offensive and defensive workflows, which means security programmes should expect more frequent low-cost exploration rather than fewer, higher-skill attempts. The practical response is to reduce the time credentials remain usable and to treat delegated access as a short-lived exposure surface, not a static entitlement.

AI attack economics: when model cost falls, the control that matters most is not whether a system can stop one attempt, but whether it can withstand thousands of cheap ones. That is where secrets lifecycle, runtime scoping, and telemetry quality become decisive, especially for programmes that already rely on the NHI Lifecycle Management Guide as their operating model.


For practitioners

  • Re-score offensive risk by cost per attempt Review AI-related threat models using the number of exploit attempts a cheap model can run before detection, not just the strength of a single model output. This helps security leadership understand where budget-friendly automation can change attacker economics.
  • Inventory exposed secrets and delegated access paths Map service accounts, API keys, tokens, and other credentials that a model-assisted workflow could probe from the outside. Prioritise internet-facing systems and third-party integrations where repeated automated attempts are cheapest to execute.
  • Shorten revocation and rotation cycles Reduce the usable lifetime of credentials that could be abused by repeated AI-driven testing. The more cheaply an attacker can search, the less tolerance you have for slow rotation and delayed offboarding.
  • Test detection against high-volume exploratory behaviour Validate whether monitoring can distinguish noisy automated probing from normal application use. The key question is whether your controls can catch repeated dead-end retries before an attacker reaches a reproducible exploit.

Key takeaways

  • Affordable AI models are making offensive security more accessible by lowering the cost of iterative exploit discovery.
  • The relevant measure is now progress per dollar, because cheap models that can run longer may matter more than premium models used sparingly.
  • Security teams should respond by tightening secrets lifecycle, delegated access scope, and detection for repeated exploratory behaviour.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Agentic AI Top 10 and MITRE ATLAS address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Agentic AI Top 10Agentic model misuse and tool-assisted offensive workflows are the article's core risk.
NIST AI RMFMANAGEThe article is about managing AI risk as model capability becomes cheaper and more scalable.
MITRE ATLASTA0007 , Discovery; TA0006 , Credential AccessThe attack pattern is exploratory discovery that can lead to credential abuse.
NIST CSF 2.0PR.AC-1Access control is relevant where AI-assisted workflows target credentials and delegated access.
NIST SP 800-53 Rev 5IA-5Credential lifecycle control is central where cheap automation can repeatedly test exposed access.

Map AI-enabled probing to agentic abuse patterns and test whether runtime controls limit repeated exploit attempts.


Key terms

  • Progress Per Dollar: A way of measuring how much useful security work an AI system can do for each unit of cost. In offensive security, this matters more than raw benchmark rank because attackers optimise for repeated attempts, runtime, and affordability rather than isolated peak performance.
  • Black-Box Vulnerability Discovery: Testing an application from the outside without source code or internal instrumentation. The attacker or model explores exposed behaviour, forms hypotheses, and retries until it finds a reproducible weakness, which makes iteration speed and cost a major part of the risk equation.
  • Exploratory Probing: Repeated automated attempts to learn how a target behaves, where it fails, and which paths are worth pursuing. It is often low-signal at first, but it can become highly effective when an AI model can run many retries cheaply and persist long enough to find a working path.
  • Delegated Access Surface: The set of credentials, tokens, API keys, and third-party permissions that can be used on a system’s behalf. This surface is especially sensitive when automation can probe it repeatedly, because weak scoping or slow revocation turns delegated access into a durable attack path.

What's in the full report

Xbow's full analysis covers the operational detail this post intentionally leaves for the source:

  • Benchmark figures and figure-by-figure comparisons across GLM, Muse Spark 1.1, Mythos, and GPT-5.5 in offensive security tasks
  • The exact evaluation setup used for black-box testing against vulnerable open-source applications
  • Cost and cache-efficiency discussion behind the progress-per-dollar view
  • Model-by-model observations on long-horizon task performance and recovery from dead ends

👉 Xbow's full post covers the model comparisons, benchmark context, and cost trade-offs behind the findings

Deepen your knowledge

NHI Mgmt Group covers identity security, NHI governance, and agentic AI through independent research, practitioner guides, and the NHI Foundation Level course, the industry's only accredited NHI security programme. It is designed for practitioners who need to connect identity control to real-world access risk across modern security programmes.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 11, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org