Subscribe to the Non-Human & AI Identity Journal

Notifications
Clear all

AI pentesting and code access: are your evaluation criteria keeping up?


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 15051
Topic starter  

TL;DR: AI pentesting buyers are being pushed to judge platforms on how they test, validate, and secure applications, with Aikido citing more than 1,000 AI pentests, a case study showing 13 issues found after a 120-hour manual pentest found none, and evidence that code access materially improves results. The governance issue is not tool preference but whether security teams can compare platforms on controls, not marketing.

NHIMG editorial — based on content published by Aikido: Buyer's Guide for AI Pentesting

Questions worth separating out

Q: How should security teams evaluate AI pentesting tools for enterprise use?

A: Judge them on representative coverage, reproducible proof, and reporting clarity, not on a single benchmark score.

Q: Why does code access matter so much in AI pentesting?

A: Code access matters because it reveals how secrets, authentication logic, and trust relationships are actually implemented.

Q: What do security teams get wrong when comparing pentesting tools?

A: They often compare output volume, interface polish, or feature lists instead of asking how the platform validates findings and limits unsafe access.

Practitioner guidance

  • Define code-access scope before procurement Require vendors to specify which repositories, environments, and artifacts the platform can inspect, and which sensitive paths are excluded.
  • Tie findings to remediation owners Map each high-confidence finding to a named application, identity, or platform owner so that secret exposure, authorization flaws, and privilege issues are tracked to closure.
  • Validate how the platform handles secrets Ask whether the tool can detect leaked keys, tokens, and embedded credentials without persisting them outside the test boundary, and verify retention controls for any captured evidence.

What's in the full report

Aikido's full guide covers the operational detail this post intentionally leaves for the source:

  • Practical vendor comparison criteria for AI pentesting platforms, including what to ask during evaluation.
  • Analysis of more than 1,000 AI pentests and the testing conditions that changed outcomes.
  • The anonymised case study showing why code access altered the number of issues found.
  • The red flags and misleading claims buyers should challenge before procurement.

👉 Read Aikido's buyer's guide for AI pentesting evaluation criteria and testing data →

AI pentesting and code access: are your evaluation criteria keeping up?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 14635
 

AI pentesting is becoming an access-governance problem, not just a testing problem. Once a platform can inspect code, it can also encounter secrets, tokens, and privileged service paths. That means buyers are no longer choosing only a testing engine, they are choosing how much sensitive identity material a tool may see during validation. The relevant governance question is whether platform access is limited, logged, and purpose-bound.

A question worth separating out:

Q: How can organisations tell whether AI pentesting is improving security?

A: They should look for reduced exposure over time, fewer repeat findings after fixes, and faster closure of issues tied to secrets or authorization logic. If retesting keeps surfacing the same problems, the programme is producing findings without changing the underlying control environment.

👉 Read our full editorial: AI pentesting promises more, but code access changes the risk equation



   
ReplyQuote
Share: