TL;DR: Security leaders are testing only 32% of their attack surface on average while 68% remains untested, according to Synack’s Gartner SRM 2026 commentary, which argues that AI-era adversaries now move on a much faster clock than conventional validation cycles. Human oversight remains essential because agentic systems can chain vulnerabilities, but current programmes still struggle to distinguish real adaptive capability from automated scanning and polished reporting.
At a glance
What this is: This analysis argues that AI is collapsing the access curve for offensive testing while enterprise pentesting still leaves most attack surface untested.
Why it matters: It matters because IAM, PAM, and broader security teams now need continuous validation, sharper prioritisation, and clearer governance over where agentic automation is useful and where human judgement still has to decide.
By the numbers:
- According to Omdia’s 2026 State of Agentic AI in Pentesting, 95% of organisations rank pentesting as a top or high priority, but they test only 32% of their attack surface each year.
- 55% of security leaders say traditional testing fails to communicate findings in a way their teams can act on.
- 64% of security leaders prefer agent-led testing with human oversight because AI cannot yet reliably chain vulnerabilities across systems.
👉 Read Synack's analysis of machine-speed adversaries and pentest coverage gaps
Context
Agentic AI is changing the economics of offensive security by compressing the time and skill needed to find exploitable paths. The core governance problem is not that testing has disappeared, but that annual and quarterly validation cycles now leave long exposure windows while machine-speed adversaries can probe far more continuously than human teams.
For IAM and PAM programmes, the real issue is control verification. If an environment cannot continuously test privilege boundaries, trust relationships, and exposed attack paths, it will keep discovering weaknesses after attackers already have a practical route in. The article’s stance reflects a broader pattern in security operations: speed alone does not equal coverage, and coverage alone does not equal prioritisation.
Key questions
Q: How should security teams evaluate agentic pentest tools?
A: Evaluate the full workflow, not the model alone. The important questions are whether the system has authoritative asset context, whether findings are verified before escalation, and whether outputs map cleanly to remediation owners. A tool that produces many findings but cannot prove them or route them effectively is creating noise, not security value.
Q: Why does low pentest coverage create a governance problem, not just a technical one?
A: Because incomplete coverage means leaders are making risk decisions with a false sense of validation. If only part of the attack surface is exercised, exposure windows remain untested, and the organisation cannot credibly claim that privilege boundaries, trust paths, and externally reachable assets are under control.
Q: What do security teams get wrong about AI safety testing?
A: The common mistake is treating AI safety testing as if it were just another security scan. It is not. Safety testing is about proving how a model or agent fails under pressure, while traditional security tooling is about who can access the system. Those are different governance questions and need different evidence.
Q: Who is accountable when agentic testing misses a critical path?
A: Accountability sits with the team that defined the scope, accepted the coverage gap, and approved the validation method. Governance frameworks should treat testing scope, review cadence, and evidence thresholds as explicit risk decisions, not informal preferences buried in tool selection.
Technical breakdown
Agentic AI vs automated scanning: why the distinction matters
Automated scanners match known patterns and can confirm previously understood issues quickly. Agentic AI goes further by planning, pivoting, and changing tactics when the first approach fails, which is why it can chain findings across systems rather than simply enumerate them. In practice, the difference is not cosmetic. A scanner can be fast and still miss the path that matters most, while an agentic workflow can adapt to new conditions if the environment, permissions, and constraints permit it. That is what makes the category strategically different for offensive testing and defensive validation.
Practical implication: treat adaptive testing as a distinct control class and require proof that a system can change strategy, not just repeat checks.
Coverage gap, not tool count, is the structural testing problem
The article’s central operational issue is coverage. Organisations may own multiple testing tools and still validate only a fraction of exposed systems because attack surfaces expand faster than assessment cycles. The challenge is compounded by heterogeneous environments, cloud sprawl, and the fact that many findings are low-context unless linked to business criticality or exploitability. This is where security validation becomes a governance problem, not just a tooling problem. Without continuous scope management and risk-based prioritisation, teams produce more findings but not necessarily more decisions.
Practical implication: measure percentage of attack surface validated, not how many scanners you operate.
Human plus AI pentesting as a control model for modern attack paths
The human plus AI model is best understood as a division of labour. AI can handle recon, enumeration, and pattern-based discovery at machine speed, while humans remain better at business logic abuse, cross-system chaining, and contextual judgement. That model mirrors other security disciplines where automation expands coverage but does not replace interpretation. For IAM-adjacent risk, this matters because privilege pathways and trust relationships often become visible only when testing crosses system boundaries. The architecture is therefore hybrid by design, not transitional by accident.
Practical implication: design pentest programmes so AI expands discovery and humans validate the exploit chain that creates real risk.
Threat narrative
Attacker objective: The objective is to reach actionable compromise through a fast, chained path that defenders did not validate in time.
- Entry occurs when machine-speed tooling identifies exposed assets, weak interfaces, or overlooked trust paths faster than annual review cycles can catch them.
- Escalation follows when the tester, or attacker, chains individual weaknesses across systems and privilege boundaries instead of treating them as isolated findings.
- Impact is realised when the resulting path reaches high-value access, sensitive data, or operational disruption before defenders have closed the gap.
NHI Mgmt Group analysis
Machine-speed adversaries expose a validation problem, not just a tooling problem. The article is right to frame the issue around coverage because organisations do not have a visibility deficit alone, they have a time-to-validate deficit. Continuous validation matters more than periodic assurance when attack paths can be discovered and chained in hours, not quarters. For identity teams, that means privilege boundaries and trust assumptions need to be tested as living controls, not annual artefacts.
Adaptive testing is becoming a governance requirement for modern attack surfaces. The useful question is no longer whether a platform can scan faster, but whether it can reason when the first path fails. That is especially relevant in environments where IAM, PAM, and service-account trust create hidden lateral movement routes. The named concept here is validation lag: the delay between exposure and meaningful control verification that leaves defenders reacting after an exploitable route already exists. Practitioners should manage for validation lag explicitly.
Human oversight remains the differentiator where business logic and multi-system chaining matter. Pure automation is still weak where the path depends on context, exception handling, or implicit trust relationships. That limitation is not a weakness to hide, it is the reason hybrid testing models will persist. Security leaders should assume AI will widen discovery and humans will continue to arbitrate which findings represent actual compromise potential.
The market is shifting from point tools toward evidence-producing security validation. Buyers will increasingly ask whether a system can demonstrate exploitability, not just enumerate issues. That changes procurement, reporting, and board communication because teams will need evidence that links technical findings to business risk. Practitioners should expect validation quality to become a primary buying criterion across offensive testing and adjacent identity governance workflows.
Identity governance is implicated whenever agentic systems touch privilege and delegation. If a testing system can reason across systems, then the same pattern will pressure any control model that assumes static access paths or single-point reviews. That is why IAM and PAM teams should pay attention even when the article reads as a pentesting story. The practical conclusion is that privilege review, access scope, and offboarding logic must be designed for continuous verification, not periodic paperwork.
What this signals
The practical signal for security leaders is that validation quality will start to matter as much as validation frequency. Programmes that cannot prove adaptive coverage will struggle to distinguish real exposure from cosmetic assurance, especially when machine-speed testing becomes normalised. Identity teams should expect tighter linkage between privilege review, trust boundaries, and continuous validation.
Validation lag: the time between an exposure appearing and a control proving it is contained or irrelevant. That delay becomes a material risk factor in any environment with delegated access, service accounts, or agent-driven workflows, and it pushes practitioners toward continuous verification models rather than cyclical reviews.
Security teams should also expect better prioritisation to replace broad alert volume as the real benchmark. Once coverage expands, the question changes from how many issues were found to which issues actually enable compromise. That is where frameworks such as the NIST AI Risk Management Framework and broader validation disciplines become operationally useful.
For practitioners
- Define validation coverage as a measurable control. Track what percentage of externally reachable assets, privileged paths, and high-value systems are actually exercised in a quarter, then tie that to risk acceptance reporting rather than tool counts.
- Require adaptive proof in vendor evaluations. Ask every testing vendor to show how the system changes tactics after a failed first attempt, because repeated pattern matching is not the same as agentic reasoning.
- Prioritise exploitable paths with business context. Use exploitability indicators, asset criticality, and privilege reach to rank findings, instead of relying on raw severity scores that overstate low-context issues.
- Test identity and trust boundaries continuously. Include service accounts, delegated access, and cross-system trust relationships in recurring validation runs so hidden lateral movement paths are surfaced before attackers find them.
Key takeaways
- The article’s core warning is that AI is collapsing offensive testing timelines while most enterprises still validate only a fraction of their attack surface.
- The evidence points to a governance gap as much as a tooling gap, because teams need adaptive validation, not just faster scanning.
- Practitioners should measure coverage, demand proof of strategy adaptation, and keep humans in the loop wherever context and chaining determine real risk.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0, NIST AI RMF, NIST SP 800-53 Rev 5 and CIS Controls v8 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | PR.AC-4 | The article centres on controlling access paths and privilege boundaries. |
| NIST AI RMF | MANAGE | AI testing governance needs lifecycle oversight, accountability, and risk treatment. |
| MITRE ATT&CK | TA0007 , Discovery; TA0008 , Lateral Movement | Adaptive testing and chained compromise map directly to discovery and movement techniques. |
| NIST SP 800-53 Rev 5 | CA-7 | Continuous control monitoring fits the article’s emphasis on ongoing validation. |
| CIS Controls v8 | CIS-7 , Continuous Vulnerability Management | The piece argues for continuous validation of exposed systems and paths. |
Align test scenarios to discovery and lateral movement behaviours rather than isolated vulnerabilities.
Key terms
- Agentic AI: Autonomous AI systems capable of planning, deciding, and taking actions — including calling APIs, writing code, and orchestrating other agents — with minimal human oversight. Agentic AI introduces new NHI risks as agents must authenticate to external services.
- Validation lag: The delay between an exposure appearing and a control proving that the exposure is understood, contained, or not exploitable. In fast-moving environments, this gap becomes a governance risk because attackers can act inside the window before defenders complete their review.
- Attack Surface Coverage: Attack surface coverage is the share of a target system's reachable components that a test meaningfully examines. It is not just enumeration of assets. It reflects whether the testing process actually reaches the endpoints, workflows, and identities most likely to contain exploitable weakness.
- Privilege Boundary: A privilege boundary is the control line that separates ordinary user actions from elevated administrative actions. When the boundary is poorly enforced, attackers can repurpose normal tools or policy logic to cross into root-level execution without going through intended approval or validation steps.
What's in the full article
Synack's full article covers the operational detail this post intentionally leaves for the source:
- How the team distinguishes automation, generative AI, and agentic AI in offensive testing workflows
- The practical breakdown of human plus AI pentesting roles, including what Sara handles versus what researchers handle
- The Omdia survey context behind the 32% coverage figure and the 55% communication gap
- The questions leaders should use to pressure-test claims about adaptive attack chaining and validation quality
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It is designed for practitioners who need to connect identity controls to broader security validation and risk management.
Published by the NHIMG editorial team on August 18, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org