TL;DR: Across two open-source applications at the same $4,000 tier, Doyensec’s benchmark found Aikido identified 49 verified vulnerabilities versus 31 for XBOW, according to Aikido. The practical issue is not only detection quality but operational friction, which now shapes how security teams evaluate AI-assisted testing.
At a glance
What this is: This is a benchmark analysis of two AI pentesting tools that found the same price tier did not translate into the same level of verified vulnerability coverage or operational effort.
Why it matters: It matters to IAM and security teams because tool choice affects how quickly vulnerabilities are found, validated, and retested, which in turn shapes risk reduction across application, identity, and access control layers.
By the numbers:
- Doyensec independently benchmarked Aikido and XBOW at the same $4,000 price tier across two real open-source applications, selected at random from 442.
- Aikido found 49 verified vulnerabilities.
- XBOW includes one retest within 30 days.
👉 Read Aikido's benchmark report on AI pentesting coverage and workflow friction
Context
AI-assisted pentesting is increasingly being evaluated as a workflow decision, not just a feature comparison. The real governance question is whether these systems reduce the time and effort required to find and verify weaknesses, or simply add another layer of tooling around the same bottlenecks in application security.
That matters to identity and access programmes because application weaknesses often create paths into secrets, tokens, privileged service accounts, and delegated access chains. When a testing workflow is slow, support-heavy, or inconsistent, exposure windows stay open longer and remediation priorities become harder to defend.
This benchmark is useful because it compares two tools at the same price tier on the same applications and then measures what practitioners actually experience during setup, validation, and retesting. That starting point is typical of real buying decisions, not an edge case.
Key questions
Q: How should security teams evaluate AI pentesting tools for enterprise use?
A: Judge them on representative coverage, reproducible proof, and reporting clarity, not on a single benchmark score. A useful tool must handle authenticated flows, multiple services, and realistic business logic, then show what it tested and why a finding is credible. If it cannot do that consistently, it is a research aid, not an enterprise control.
Q: Why does setup friction matter in security testing programmes?
A: Because delay before the first scan delays discovery, reporting, and remediation. When onboarding requires contracts, repeated emails, or scan restarts, the testing programme becomes slower and less scalable. That matters most where the findings can expose secrets, credentials, or access paths that need quick verification and closure.
Q: What breaks when retesting is too limited?
A: Remediation confidence drops because teams cannot verify fixes on their normal schedule. A narrow retest window forces rushed validation or leaves issues open until the next cycle, which weakens evidence for risk closure. For teams running continuous assurance, retesting needs to fit the pace of actual change, not the vendor’s default limit.
Q: What does a benchmark like this reveal about AI-assisted security testing?
A: It shows that real value comes from the full control loop, not just the scan result. Coverage, turnaround time, and retest ease all shape whether findings are useful enough to drive action. Practitioners should treat AI testing as an operating model choice, not a feature checkbox.
Technical breakdown
What independent AI pentesting benchmarks actually measure
A useful benchmark does more than count findings. It should compare verified vulnerabilities, severity distribution, overlap between tools, and the operational steps required to obtain results. In this case, the report distinguishes detection volume from usefulness by showing that both tools found real issues but differed materially in coverage and workflow friction. That distinction matters because security teams do not buy findings in isolation. They buy the ability to convert those findings into remediation work, retesting, and reporting without introducing extra coordination overhead.
Practical implication: evaluate AI pentesting tools on verified coverage, retest flow, and delivery friction, not just headline vulnerability counts.
Why setup friction changes the economics of security testing
Setup time is part of control effectiveness because delayed onboarding delays discovery. If one workflow needs contracts, support emails, and scan restarts while another begins in minutes, the second path creates a shorter time to insight and a lower burden on the team consuming the results. For identity-heavy environments, that also means faster visibility into exposed tokens, misused sessions, and application logic flaws that can lead into broader access paths. The technical issue is not only tool speed, but the amount of human orchestration required before testing even begins.
Practical implication: measure time-to-first-scan and escalation overhead as governance metrics, especially where app findings can intersect with secrets and access abuse.
Why retest design affects remediation confidence
Retesting is part of the control loop, not an afterthought. Unlimited or low-friction retesting supports rapid verification of fixes, while a narrow retest window forces teams to bundle verification into a short period and can leave remediation evidence incomplete. In practice, this matters when findings touch authentication flows, API permissions, or exposed credentials, because teams need to prove that the issue was removed, not merely queued for the next cycle. A benchmark that includes retest constraints is therefore closer to real operations than a simple scan comparison.
Practical implication: align retesting terms with your remediation cadence so verification does not become a separate operational bottleneck.
NHI Mgmt Group analysis
AI pentesting should be judged as an operational control, not a demo capability. The benchmark shows that output quality alone is not enough to explain tool value. Coverage, turnaround time, and retest accessibility determine whether findings become usable security signal or just another backlog item. For practitioners, that means the real question is how quickly a tool improves risk decisions across the application and identity surfaces it exposes.
Coverage drift is the name for the gap between findings volume and actionable security value. Two tools can report real vulnerabilities and still produce very different governance outcomes if one finds broader issue sets and returns them faster. That is especially relevant where application weaknesses can expose secrets, access tokens, or privileged workflows. Practitioners should treat coverage drift as a procurement and assurance concern, not just a technical benchmarking nuance.
AI-assisted testing will increasingly be evaluated through the lens of control-loop speed. The faster a tool surfaces verified weaknesses and supports retesting, the more it can fit into continuous assurance programmes. Slow onboarding, contract friction, and repeated restarts all weaken that loop. For security leaders, the signal is clear: adoption decisions should prioritise the shortest path from discovery to validated remediation.
This benchmark also shows that workflow friction is part of attack surface reduction. If a tool takes 11 days to produce output after engagement start, the enterprise loses time that could have been spent closing exposure. That is not just an efficiency issue. It affects how quickly teams can reduce exploitable conditions across web applications, APIs, and any downstream identity dependencies they reveal. The practitioner conclusion is to treat latency as a governance metric.
For identity programmes, application testing remains a front door to credential and privilege risk. Even when the article is about pentesting coverage, the downstream value often sits in discovering how apps expose sessions, tokens, secrets, and access pathways. That makes the security architecture around AI pentesting relevant to IAM, PAM, and NHI governance. Teams should use benchmark results to prioritise the tooling that most quickly reveals access-adjacent weaknesses.
What this signals
Coverage drift will become a more important procurement signal as AI-assisted testing matures. Security teams should ask whether a tool produces broader verified coverage, faster validation, and lower coordination overhead, because those factors determine whether findings change behaviour or simply increase noise. The relevant benchmark is not the scan itself but the speed at which it moves a team from discovery to closure.
Where AI testing intersects with identity risk, the question becomes whether application findings are being pushed into secrets, token, and access review workflows quickly enough. That is where the The State of Secrets in AppSec findings matter: remediation gaps are measured in weeks, not hours, so slow testing pipelines compound exposure. Practitioners should therefore connect application testing output to identity governance queues and incident prioritisation.
For security programmes, the next step is to treat benchmark evidence as a governance input. If one workflow can identify and validate problems in minutes while another needs repeated support intervention, the programme should value the former for operational assurance and not just offensive testing. That shift matters most in environments where application flaws can turn into credential exposure, delegated access abuse, or privileged path creation.
For practitioners
- Benchmark against verified coverage, not raw findings counts Compare tools on unique verified vulnerabilities, severity distribution, and overlap across the same applications before procurement decisions. Use a fixed sample set so the benchmark reflects repeatable coverage rather than marketing claims.
- Measure onboarding friction as part of tool evaluation Track time to first scan, number of support exchanges, contract steps, and restarts required before results are produced. If the workflow needs repeated human intervention, treat that as a control cost, not a procurement detail.
- Align retest terms with remediation cadence Require retesting terms that match how quickly your team closes findings, especially for issues that can expose secrets or access paths. A short retest window can leave validation work incomplete and slow evidence-based closure.
- Feed application findings into identity review workflows Escalate issues that expose tokens, API keys, privileged sessions, or delegated access into IAM and PAM review queues, not just application tickets. That keeps identity risk visible when a web flaw turns into an access problem.
Key takeaways
- The benchmark suggests that AI pentesting value is defined by verified coverage and operational friction, not just by the number of findings returned.
- When a tool takes days to deliver output or needs repeated manual support, its security value is reduced even if its findings are real.
- Teams should evaluate AI testing through remediation speed, retest flexibility, and the downstream identity risk exposed by application weaknesses.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0, NIST SP 800-53 Rev 5 and NIST AI RMF set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | DE.CM-7 | The article evaluates how quickly security tooling detects verified weaknesses. |
| NIST SP 800-53 Rev 5 | RA-5 | Vulnerability scanning and analysis maps directly to this control family. |
| NIST AI RMF | MEASURE | The benchmark is fundamentally a measurement exercise for AI-enabled security tooling. |
| MITRE ATT&CK | TA0007 , Discovery; TA0040 , Impact | The findings help identify weaknesses before adversaries can discover and exploit them. |
Validate that AI testing results feed RA-5 workflows for prioritisation and remediation tracking.
Key terms
- Verified vulnerability: A verified vulnerability is a weakness that has been manually confirmed rather than inferred by automation alone. In benchmark testing, this matters because it filters out noisy results and makes coverage comparisons more meaningful for remediation planning and risk reporting.
- Coverage Drift: The gap between a security policy that exists on paper and the parts of the environment where it is actually enforced. In identity programmes, coverage drift appears when exceptions, legacy apps, or bypass paths allow controls like MFA to be selectively ignored.
- Retest window: A retest window is the period in which a security team can re-run validation after fixes are applied. Short windows can force rushed verification and reduce confidence that a remediation actually worked, especially when the affected system changes frequently.
What's in the full report
Aikido's full benchmark report covers the operational detail this post intentionally leaves for the source:
- Side-by-side evidence tables showing the exact verified findings in Fider and Photoview.
- Workflow detail on the 20-minute setup path, including what was required to start scanning.
- Benchmark notes on retest scope, support overhead, and infrastructure interruptions during the XBOW engagement.
- Full severity and false-positive breakdowns for each tool across the two applications.
Deepen your knowledge
NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It gives practitioners a stronger basis for connecting application findings to access, credential, and lifecycle controls.
Published by the NHIMG editorial team on August 2, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org