Subscribe to the Non-Human & AI Identity Journal

Notifications
Clear all

AI vulnerability exploitation models: what cost-efficient accuracy really means


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 13011
Topic starter  

TL;DR: GLM-5.2 exploited 67% of planted vulnerabilities at $1.13 per lead, making it the cheapest configuration in its benchmark, according to TENZAI. Cost efficiency is improving faster than dependable coverage, so security teams should treat AI offensive tooling as a scaling factor, not a substitute for control depth.

NHIMG editorial — based on content published by TENZAI: GLM-5.2 is the most cost-efficient cyber model we've tested, but it is not #1

Questions worth separating out

Q: How should security teams evaluate AI claims in cybersecurity tools?

A: They should evaluate the tool by its actual decision behaviour, not by marketing language.

Q: Why does cost matter in AI-assisted vulnerability exploitation?

A: Cost determines whether a model can be used continuously or only in limited bursts.

Q: What should teams get wrong about AI exploitation benchmarks?

A: They often treat a benchmark score as proof of real-world readiness.

Practitioner guidance

  • Benchmark AI offensive tools against exploit success and cost Measure models on the percentage of planted vulnerabilities they can actually exploit, alongside cost per lead, so procurement choices reflect operational value rather than headline capability.
  • Separate subagent responsibilities in AI security workflows Require different controls for lead generation, proof-of-concept creation, and execution so a single workflow cannot silently move from analysis to active exploitation.
  • Prioritise findings that change privilege boundaries Triage AI-discovered issues by whether they expose credentials, enable privilege escalation, or open lateral movement paths, using access impact as the first filter.

What's in the full report

TENZAI's full research covers the benchmark mechanics and model-by-model comparisons this post intentionally leaves at summary level:

  • Per-model recall curves across GLM-5.2, GPT-5.5, and Opus-4.8 at different reasoning levels
  • Cost-per-vulnerability comparisons that show where higher reasoning effort changes the trade-off
  • Benchmark setup details for the planted vulnerabilities and enterprise application scenarios
  • The role of Bonzai and the subagent workflow in turning leads into exploit proof-of-concepts

👉 Read TENZAI's benchmark analysis of GLM-5.2, GPT-5.5, and Opus-4.8 →

AI vulnerability exploitation models: what cost-efficient accuracy really means?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 12595
 

Cost efficiency is becoming a governance issue, not just a procurement one. A model that can exploit 67% of planted vulnerabilities at low cost changes how teams scale offensive validation, because cheaper runs mean more frequent testing and wider coverage. That helps security programmes, but it also lowers the barrier to misuse by less sophisticated actors. The practical conclusion is that cost per exploit will become part of risk assessment, not just a line item in tooling selection.

A question worth separating out:

Q: How do identity controls change when AI tools can generate exploit paths?

A: Identity controls become the blast-radius limiters after a vulnerability is found. Even when AI uncovers a weak point, least privilege, secrets hygiene, and privileged session controls determine whether the issue becomes a local defect or an incident with lateral movement and credential exposure. That makes IAM and PAM part of exploit containment.

👉 Read our full editorial: AI vulnerability exploitation models are cheaper, but still miss depth



   
ReplyQuote
Share: