Join our Newsletter — 33% off our NHI Course

Notifications
Clear all

NexBench and long-running agent evals: what do teams need to know?


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 20360
Topic starter  

TL;DR: Existing cybersecurity evals overstate agent performance because they are too boxed in, while real environments require long-running, cost-aware exploit discovery across messy, multi-stage targets, according to MindFort’s NexBench benchmark. The practical lesson is that offensive security for AI agents now needs validation, runtime, and economic realism, not just benchmark scores.

NHIMG editorial — based on content published by MindFort: Introducing NexBench, MindFort's internal model evaluation for offensive security agents

Questions worth separating out

Q: How should teams evaluate AI offensive security agents in realistic environments?

A: Teams should test agents in stateful environments that allow chained vulnerabilities, repeated runs, and independent validation.

Q: Why do long-running AI agents create a different security governance problem?

A: Long-running agents can preserve context, adapt to failures, and continue probing until they find a path forward, which makes them closer to real attackers than short-lived benchmark runs.

Q: What do security teams get wrong about using benchmark scores to judge AI coding risk?

A: They often treat one benchmark number as proof of broad security quality.

Practitioner guidance

  • Require validated exploit reproduction Use judge-based re-execution, not single-pass model output, before accepting an agent finding as real.
  • Measure runtime coherence explicitly Track how long an agent can preserve reasoning across multi-stage tasks, especially when attacks require several hours of continuous execution.
  • Add cost-per-finding thresholds Set a ceiling for token spend and compute per validated finding so a strong model that is economically unusable does not enter production testing pipelines.

What's in the full report

MindFort's full blog covers the operational detail this post intentionally leaves for the source:

  • The exact benchmark setup and scoring logic used to compare models across low, medium, and high findings.
  • The model-by-model leaderboard, including runtime, token use, and cost-per-finding calculations.
  • The specific efficiency trade-offs between hosted and local models in offensive security workflows.
  • The benchmark methodology for judge-agent validation and exploit re-execution.

👉 Read MindFort's analysis of NexBench and long-running offensive AI evals →

NexBench and long-running agent evals: what do teams need to know?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 4 months ago
Posts: 19951
 

Benchmark realism is now a governance issue, not just a research preference. Offensive AI evaluations that ignore chained vulnerabilities, stateful access, and long runtime can produce misleading confidence. In AI security, the question is no longer whether a model can solve a task in isolation, but whether it can sustain reasoning across the conditions that real attackers exploit. The right lens is operational fidelity, especially when AI agents may be used for red teaming or autonomous security workflows.

A question worth separating out:

Q: How can organisations decide whether AI-assisted penetration testing is worth using?

A: Start by comparing validated findings per dollar, not raw model quality. Then add runtime ceilings, reproducibility checks, and environment realism to the decision. If a model is cheap but cannot sustain execution or confirm results, it is not ready for operational security workflows.

👉 Read our full editorial: NexBench shows model evals still miss real offensive security conditions



   
ReplyQuote
Share: