TL;DR: 64% of organisations prefer agent-led pentesting with human oversight, while 87% report high or complete trust in agentic AI and 69% require at least 85% accuracy before using it in production, according to Synack’s commissioned Omdia study. The data suggests security teams are treating explainability, guardrails, and human validation as governance requirements, not optional features.
At a glance
What this is: This is Synack’s analysis of how enterprises are adopting agentic AI for pentesting, with human oversight emerging as the preferred operating model and accuracy the main trust threshold.
Why it matters: It matters because agentic AI in security testing raises governance questions about validation, scope, and accountability that IAM, PAM, and security architecture teams will need to formalise across both human and machine-led workflows.
By the numbers:
- 64% of organisations identify agent-led with human oversight as their preferred operational model.
- 87% of organizations have moved beyond the evaluation phase and are actively planning, piloting, or using agentic AI for pentesting.
- 69% of those surveyed said they require an accuracy level of at least 85% compared with manual testing.
- 95% of organizations anticipate that agentic AI will displace traditional pentesting services to some degree.
👉 Read Synack's analysis of human-validated agentic AI pentesting
Context
Agentic AI is changing pentesting from a purely human exercise into a mixed operating model where software can plan, test, and validate at machine speed. The governance challenge is no longer whether AI can assist security testing, but how teams bound its scope, verify its outputs, and decide when human judgement must remain mandatory. That question is now directly relevant to security leaders, IAM teams, and those responsible for identity-adjacent controls such as access validation and privilege review.
Synack’s research suggests the market is converging on a supervised model rather than full autonomy. That is consistent with how security teams usually adopt new automation: they want speed and scale, but they do not want to give up evidential certainty, explainability, or accountability. The operational issue for practitioners is how to preserve trust when the testing engine itself becomes an active decision-maker.
For teams already using AI in security workflows, the article reflects a typical adoption pattern: move from pilot to production only after guardrails, verification, and human review are in place. The starting point is increasingly typical, but the insistence on oversight is the more important signal.
Key questions
Q: How should security teams govern agentic pentesting tools in production-like environments?
A: Treat them as delegated systems with explicit scope, named ownership, and approval checkpoints. The control objective is not to stop automation, but to ensure that any step affecting production stability, compliance, or rules of engagement requires a human decision before execution continues.
Q: Why do organisations keep human oversight in agentic security testing?
A: Human oversight remains necessary because security testing is not only about speed. Teams need contextual judgement, safe handling of destructive actions, and evidence they can defend to auditors and business owners. The more advanced the automation, the more important it becomes to verify what the system concluded and why.
Q: What breaks when agentic AI testing is allowed to run without strong guardrails?
A: Without guardrails, an AI testing system can exceed scope, use unsafe commands, or generate findings that cannot be trusted. That creates operational risk, inflated remediation queues, and loss of confidence in the whole programme. The failure is not just technical. It is governance failure around authority and containment.
Q: Who is accountable when an AI system used for security testing crosses into abuse?
A: Accountability sits with the organisation that grants access, defines scope, and approves the workflow. That usually includes security leadership, platform owners, and the teams managing the AI toolchain. If a model can act on behalf of a business process, the business must control the identity, permissions, and audit trail behind it.
Technical breakdown
Agent-led pentesting versus fully autonomous testing
Agent-led pentesting means AI systems coordinate reconnaissance, test execution, and preliminary validation, while humans remain responsible for scope, interpretation, and final acceptance. Fully autonomous testing removes that human checkpoint, which increases scale but also increases the risk of unsafe actions, false confidence, and evidence that cannot be defended to stakeholders. In practice, the difference is not just speed. It is the control model governing who can authorise actions, when findings become credible, and how destructive behaviour is prevented.
Practical implication: treat agentic pentesting as a governed workflow with approval, scope, and review controls, not as a replacement for human-led validation.
Why accuracy and explainability matter in security testing
Accuracy in agentic pentesting is not just about finding vulnerabilities. It is about confirming exploitability, suppressing noise, and producing results that can survive operational scrutiny. Explainability matters because security teams need to understand why the system reached a finding, what evidence it used, and whether the same conclusion would hold in a different environment. Without that, the output becomes another source of triage burden rather than a decision aid.
Practical implication: require proof-based validation and decision traces before allowing AI-generated findings into remediation or reporting workflows.
Guardrails are the control layer that keeps AI within test scope
Guardrails are the rules that stop an agentic system from acting outside approved assets, commands, and objectives. In pentesting, that includes scope enforcement, destructive-command blocklists, and constraints on where the system may probe. This is an identity and authorisation problem as much as a testing problem, because the agent must be issued limited rights and monitored like any other privileged runtime actor. If the guardrails fail, the testing platform can become a source of production risk.
Practical implication: define the agent’s allowed actions, asset boundaries, and escalation rules before deployment, and review them as you would privileged access.
Threat narrative
Attacker objective: The objective is to turn machine speed into actionable exploit discovery without losing evidential certainty or control over the environment being tested.
- Entry begins when an AI testing system is granted access to in-scope assets and uses those permissions to map the environment and probe for weaknesses.
- Escalation occurs if the system is allowed to attempt exploits or execute destructive commands without strict guardrails and human verification.
- Impact is created when false positives, unsafe actions, or unverified findings distort remediation decisions or cause operational disruption.
NHI Mgmt Group analysis
Human oversight is becoming the control boundary for agentic security testing. The article shows that organisations are not adopting AI testing as an autonomy problem, they are adopting it as a supervised operations problem. That is the right instinct. In security testing, the value of automation collapses if the output cannot be defended, explained, and independently verified. Practitioners should treat the oversight layer as part of the control plane, not a temporary concession.
Verification trust gap: the real issue is not whether AI can act, but whether its results are trustworthy enough to drive remediation. Synack’s data highlights a familiar pattern in security tooling adoption: teams tolerate automation only when it produces evidence they can validate. That is especially true in identity-adjacent testing, where access scope, exploitability, and business context determine whether a finding matters. The governance question is whether your process can distinguish proof from noise at production pace.
Agentic AI pentesting creates a new class of privileged runtime actor. Once an AI system can plan actions and execute tests, it needs tightly bounded permissions, approval checkpoints, and logging comparable to privileged human access. That brings IAM and PAM concepts into a domain that historically treated tooling as inert. Security teams should assess agent identity, not just model quality, because runtime authority is now part of the risk surface.
Human-validated AI is more likely to expand than replace security expertise. The article’s own findings suggest that the most experienced users are not moving toward unchecked autonomy. They are moving toward a model where AI covers scale and humans cover judgment, which is closer to a control framework than a labour replacement story. The practitioner conclusion is straightforward: build governance for collaboration between machine execution and human accountability, not for a fantasy of hands-off security testing.
What this signals
Verification trust gap: security teams will increasingly judge AI tools by the quality of their evidence, not by how much work they automate. That shifts procurement and governance conversations toward explainability, scope control, and auditability, especially where testing output feeds risk acceptance or remediation.
Agentic AI in security testing is also a reminder that runtime authority now matters for non-human systems. If the agent can act, then its permissions, logging, and containment need to be managed with the same seriousness applied to privileged human access, including alignment with the NIST AI Risk Management Framework.
Teams should expect more demand for governed collaboration between machine execution and human review. The practical signal is not full autonomy, but a higher bar for evidence and an insistence that AI-driven findings can survive challenge from security, audit, and operations stakeholders.
For practitioners
- Define the agent’s scope as a privileged access policy Document which assets, test types, and commands an AI pentesting system may use, and require explicit approval for anything outside that boundary. Treat the agent like a privileged runtime with least-privilege access, not like an ordinary software utility.
- Require proof-based validation before findings reach remediation Only allow findings into ticketing or reporting workflows when they include independent re-test evidence, exploit confirmation, and a clear explanation of why the result is credible. That reduces noise and prevents teams from acting on unverified output.
- Add human sign-off for destructive or high-impact actions Block commands or actions that could delete data, alter systems, or disrupt production unless a human reviewer explicitly approves the test path. This is especially important where AI systems interact with live environments rather than isolated labs.
- Log agent decisions as audit evidence Capture the agent’s inputs, target selection, exploit attempts, and verification steps so security leaders can reconstruct why a finding was produced. That evidence supports both remediation and accountability when stakeholders question the result.
Key takeaways
- Agentic AI pentesting is being adopted as a supervised operating model, not a hands-off one.
- Trust in AI security testing depends on explainability, proof, and guardrails more than raw speed.
- Security teams should govern AI testing systems like privileged actors with bounded authority and human oversight.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and MITRE ATT&CK address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | Agentic AI testing raises guardrails, scope, and verification issues covered by agentic application guidance. | |
| NIST AI RMF | GOVERN | Human oversight, accountability, and auditability are core AI governance concerns in this article. |
| NIST CSF 2.0 | PR.AC-4 | Agent permissions and scope enforcement align with least-privilege access governance. |
| NIST SP 800-53 Rev 5 | AC-6 | Least privilege is directly relevant to bounding what the agent may access or execute. |
| MITRE ATT&CK | TA0006 , Credential Access; TA0004 , Privilege Escalation | The article discusses exploit execution and verification against real environments, which maps to attacker tactics. |
Map agent permissions, tool use, and output verification to the agentic AI risk areas that affect production testing.
Key terms
- Agent-Led Pentesting: A testing model where AI systems coordinate discovery and exploit attempts while humans retain oversight and final validation. It combines machine scale with human judgment, making governance and evidence quality central to whether the output is operationally trustworthy.
- Human Oversight: Human oversight is the requirement that a person remains responsible for reviewing, approving, or correcting AI-driven output before it causes a material action. In governance terms, it is the control that prevents automation from becoming unowned authority.
- Activation Trust Gap: The activation trust gap is the difference between trusting data because it is protected and governing it because it is being reused. It appears when organisations move data from backup or archival systems into AI pipelines without reapplying access, sensitivity, and consumer controls.
- Guardrails: Guardrails are policy controls that inspect prompts and model outputs against defined safety, privacy, and compliance rules. In AI operations, they reduce harmful language and disclosure risk, but they do not replace entitlement management, logging, or identity governance for the systems that call the model.
What's in the full report
Synack's full research covers the operational detail this post intentionally leaves for the source:
- Survey methodology and respondent mix behind the 64% oversight preference and 87% adoption figures
- Breakdown of how production users define acceptable accuracy, transparency, and guardrail thresholds
- Examples of the agent-led testing model, including verification workflow and human review steps
- FAQ-level detail on Sara Pentest and Sara Triage capabilities for teams evaluating implementation
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It is designed for practitioners who need to manage identity risk across human, non-human, and emerging AI-driven workflows.
Published by the NHIMG editorial team on August 19, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org