Join our Newsletter — 33% off our NHI Course

Notifications
Clear all

Application pentesting and AI agents: where human researchers still matter


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 20026
Topic starter  

TL;DR: Application pentesting is becoming unusually amenable to AI because it offers fast, objective feedback loops, and FireCompass says its agents now beat top researchers 60 to 70% of the time with under 2% false positives. The governance question is no longer whether AI can assist testing, but which parts of validation-heavy security work should remain under human direction.

NHIMG editorial — based on content published by FireCompass: Why AI May Disrupt Application Pentesting Earlier Than Most Security Teams Expect

Questions worth separating out

Q: How should security teams implement AI penetration testing for agents and models?

A: Start with the highest-risk workflows first, especially agents that can access SaaS data, APIs, or approval paths.

Q: Why do agentic pentest systems improve faster than traditional tools?

A: They improve quickly when each step creates immediate feedback.

Q: What do security teams get wrong about AI-generated penetration testing findings?

A: The main mistake is treating AI output as proof rather than as a lead.

Practitioner guidance

  • Define which pentest steps can be machine-verified Separate discovery, validation, and interpretation into different control layers so AI handles only the parts with objective evidence and stable success criteria.
  • Retire benchmark-only evaluation for offensive agents Use live applications, stateful workflows, and researcher-versus-agent testing instead of relying on public benches that can saturate quickly.
  • Tie agent output to IAM and application control owners Route findings on authorization failures, privilege boundaries, and trust-boundary breaks to the teams that own those controls, not only to red-team or AppSec operators.

What's in the full article

FireCompass's full blog covers the operational detail this post intentionally leaves for the source:

  • A more granular breakdown of the researcher-versus-agent evaluation method used once public benchmarks saturated
  • The progression of agent versions and what changed in tooling, memory, and validation as performance improved
  • Specific examples of the control-validation tasks where agents outperformed human researchers
  • The operational framing behind under 2% false positives and why that threshold matters for teams using offensive automation

👉 Read FireCompass's analysis of why AI may disrupt application pentesting earlier than expected →

Application pentesting and AI agents: where human researchers still matter?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 4 months ago
Posts: 19617
 

Verifiability is becoming the decisive force multiplier in offensive security. When a workflow gives an AI system rapid, objective feedback, the system can outpace humans in narrow but important tasks. Application pentesting fits that pattern unusually well because many steps are measurable, repeatable, and evidence-driven. That means the market will increasingly reward systems that can close the loop between action and validation, not just systems that can generate plausible hypotheses. Practitioners should treat verifiability as a design criterion for both attack and defense tooling.

A question worth separating out:

Q: How should organisations divide work between AI agents and human researchers?

A: Let AI agents handle repetitive discovery, validation, and retry-heavy testing, while humans set objectives, interpret ambiguous findings, and decide what matters commercially. That division keeps the machine in the high-feedback loop and preserves human judgment for risk prioritisation, context, and supervisory control over the offensive system.

👉 Read our full editorial: AI may disrupt application pentesting sooner than expected



   
ReplyQuote
Share: