Join our Newsletter — 33% off our NHI Course

Notifications
Clear all

PWNBench and web app pentesting: are current AI evals enough?


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 19382
Topic starter  

TL;DR: Frontier LLMs vary sharply in web application pentesting, with recall, precision, and cost moving very differently across models and test-time compute settings, according to Novee. The broader lesson is that cyber evals must measure end-to-end attack workflow and false positives, not just raw finding volume.

NHIMG editorial — based on content published by Novee: PWNBench-v0.1: Evaluating frontier models for web application pentesting

By the numbers:

Questions worth separating out

Q: How should security teams evaluate AI pentesting tools for enterprise use?

A: Judge them on representative coverage, reproducible proof, and reporting clarity, not on a single benchmark score.

Q: Why do AI pentesting agents need governance beyond model selection?

A: Because the model is only one part of the system.

Q: What do security teams get wrong about benchmark scores for agentic systems?

A: They often treat a benchmark result as a stable property of the system, when it is really a snapshot of behaviour under specific conditions.

Practitioner guidance

  • Define severity-weighted acceptance criteria Require pentesting or agentic security tools to report precision, recall, and severity together, with high-severity false positives tracked separately from low-severity noise.
  • Cap test-time compute by use case Set different reasoning and multi-run budgets for discovery, validation, and reporting tasks so teams can compare outcomes without uncontrolled cost drift.
  • Bind AI security tools to governed credentials Treat credentials used by web-testing agents as a managed NHI estate, with scoped access, audit logging, and revocation tied to the testing session.

What's in the full report

Novee's full blog post covers the operational detail this post intentionally leaves for the source:

  • Per-model benchmark curves for recall, precision, and cost across all 11 frontier systems.
  • The exact evaluation harness design and how Novee validated its ground-truth issue labels.
  • Severity-level splits showing how high-severity findings change the ranking.
  • The model-by-model discussion of where added reasoning effort improved or degraded performance.

👉 Read Novee's PWNBench-v0.1 analysis of frontier model pentesting performance →

PWNBench and web app pentesting: are current AI evals enough?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 18973
 

Precision is the more important control signal than raw recall. A pentesting system that finds many issues but emits too many false positives does not reduce operational risk effectively. This benchmark shows why security teams should measure usefulness, not just output volume. For AI security programmes, the real question is whether the system can support triage and validation without forcing analysts to relearn the target from scratch.

A question worth separating out:

Q: How can organisations control risk when AI systems are given pentesting credentials?

A: Scope those credentials like any other non-human identity. Give the agent only the access needed for the test, log every action, and revoke credentials as soon as the session ends. If the system can explore live applications, its access should be time-bound, observable, and easy to terminate.

👉 Read our full editorial: PWNBench shows cyber evals need precision, not recall alone



   
ReplyQuote
Share: