Join our Newsletter — 33% off our NHI Course

Notifications
Clear all

Agentic AI pentesting and SQLi labs: what should teams trust?


(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 18004
Topic starter  

TL;DR: An agentic AI pentest system solved 10 of 10 PortSwigger SQL injection labs in January 2026, with 90% completed within a standard practitioner time window and one missed by 18 seconds, according to Synack. The result shows that agentic testing can scale coverage, but human validation still matters when tool-driven systems are used to assess real attack surface.

NHIMG editorial — based on content published by Synack: Benchmarking Synack’s Agentic AI Against PortSwigger SQLi Labs

Questions worth separating out

Q: How should security teams use AI-assisted pentesting without losing control of evidence quality?

A: Use AI-assisted pentesting as a decision-support layer, not a decision authority.

Q: When does AI-assisted pentesting reduce more risk than manual testing alone?

A: It reduces more risk when environments are large, distributed, and changing faster than a traditional engagement can keep up.

Q: What do organisations get wrong about vulnerability discovery?

A: They often treat discovery as proof of risk.

Practitioner guidance

  • Define acceptance criteria for agentic pentest results Require reproducible evidence, clear exploit steps, and human validation before findings can enter remediation or audit workflows.
  • Use agentic testing to triage high-volume web estates Apply AI-assisted testing first where SQL injection and related input-handling issues are common, then reserve manual effort for the highest-risk applications.
  • Separate discovery from assurance decisions Make sure the team that receives agentic findings is not the only team validating them, so quality control remains independent.

What's in the full report

Synack's full blog covers the operational detail this post intentionally leaves for the source:

  • The lab-by-lab benchmark table for all 10 PortSwigger SQLi challenges and the exact completion times.
  • The testing approach Sara used to adapt payloads, identify database type, and handle blind or out-of-band cases.
  • The specific practitioner framing Synack uses to compare human-led and agentic-led pentesting by asset importance and speed.
  • The follow-on vulnerability classes Synack says it will benchmark next, including broken authentication and SSRF.

👉 Read Synack's benchmark of Sara Pentest against PortSwigger SQL injection labs →

Agentic AI pentesting and SQLi labs: what should teams trust?

Explore further

View Full Forum →  |  NHI Foundation Course →



   
Quote
(@mr-nhi)
Member Moderator
Joined: 3 months ago
Posts: 17593
 

Agentic AI pentesting is becoming a governance problem, not just a testing efficiency story. The value proposition is not merely faster vulnerability discovery. It is the ability to compress reconnaissance, exploitation validation, and reporting into a smaller operating window, which changes how teams prioritise remediation. That matters to IAM and security leaders because the decision to trust agentic output now affects risk governance, not just test throughput.

A question worth separating out:

Q: Should organisations replace manual pentests with agentic testing?

A: No. Agentic testing is best treated as a high-frequency validation layer that expands coverage and speed, while humans remain essential for scoping, exception handling, and adjudicating complex findings. The practical model is hybrid: automation for breadth and repeatability, humans for judgement and edge cases.

👉 Read our full editorial: Agentic AI pentesting benchmarks raise new governance questions for SQLi



   
ReplyQuote
Share: