TL;DR: Continuous web attack-surface coverage now depends on balancing automation depth with human-reviewed assurance for auditors and stakeholders, according to Terra. Riskified’s use of Terra’s agentic AI-driven pentesting shows why the governance issue is no longer whether automation can find more, but whether its outputs remain credible, controlled, and acceptable in regulated programmes.
At a glance
What this is: This is an analysis of how agentic AI can extend penetration testing depth while still preserving the human oversight needed for audit-ready assurance.
Why it matters: It matters because security teams running web applications, compliance programmes, and AI-enabled workflows need testing models that scale without losing evidence quality, control, or governance acceptability.
👉 Read terra's analysis of agentic AI pentesting and audit assurance
Context
Agentic AI is changing penetration testing by widening coverage and increasing testing frequency, but it does not remove the need for human review where auditability, safety, and reporting credibility matter. In practice, the problem is not just test depth, it is whether the testing process can still produce outputs that governance teams and auditors will trust.
For IAM and NHI practitioners, the identity angle sits in the controls around the agent itself: who authorises its actions, what systems it can reach, and how its outputs are validated before they inform remediation or assurance. That makes this a governance problem as much as a technical one, especially in programmes that already struggle with machine trust and delegated access boundaries.
Key questions
Q: How should security teams govern AI agents used for offensive testing?
A: Treat offensive AI agents as distinct workloads with explicit ownership, scoped tools, and logged approvals. Give them only the environments, credentials, and actions needed for authorised testing. Separate research targets from production systems, and review retries, data access, and output handling as part of standard governance, not as an afterthought.
Q: Why do autonomous testing tools still need human oversight?
A: Autonomous testing tools still need human oversight because auditors and governance teams need validated evidence, not just generated findings. Human review confirms exploitability, checks safety, and ensures the report can support accountability. Without that control, penetration testing may become faster but less credible, especially in regulated environments where assurance artefacts must be defensible.
Q: What breaks when pentest automation is not tied to audit controls?
A: When pentest automation is not tied to audit controls, teams often end up with findings that are difficult to trust, reproduce, or sign off. That creates a gap between technical discovery and governance acceptance. The programme may look more active, but it delivers less usable assurance and can even introduce safety risk if execution is not properly bounded.
Q: Should organisations treat agentic security tools like non-human identities?
A: Yes. Once a security tool can choose actions and execute them independently, it needs the same lifecycle discipline applied to other non-human identities. That means unique credentials, scoped permissions, monitoring, and revocation. The governance question is not whether the tool is intelligent, but whether its access is controllable and attributable.
Technical breakdown
How agentic AI expands pentest coverage without replacing analysts
Agentic AI systems can independently sequence actions, choose paths through an attack surface, and continue exploring across many application states where manual testing would stall. That makes them useful for continuous validation of large web estates, but only if their autonomy is constrained by explicit scope, safety checks, and reviewable output. In practice, the value comes from breadth and persistence, not from removing humans from the loop. The technical question is how to let the system search aggressively while still preserving evidence quality and bounded execution.
Practical implication: define hard scope limits and review checkpoints before agentic tools are allowed to run against production-facing systems.
Why audit-ready pentesting still depends on human oversight
Audit assurance requires more than a vulnerability finding. It requires reproducible evidence, reviewer accountability, and outputs that can stand up to internal control testing or external scrutiny. Fully autonomous testing can struggle here because auditors need to see that results were validated, not merely generated. Human oversight is therefore part of the control design, not a cosmetic layer. The human role shifts from manually discovering every issue to validating exploitability, confirming safety, and signing off on artefacts that meet governance expectations.
Practical implication: keep human approval in the reporting and validation path wherever pentest results feed compliance, board reporting, or formal risk acceptance.
Agentic AI in security testing creates its own identity and access boundaries
When an AI system is allowed to test infrastructure, it becomes a governed actor with permissions, tool access, and execution constraints. That makes its identity relevant in the same way service accounts and automation credentials are relevant in other programmes. If the testing agent can reach more systems than intended, or if its actions are not attributable, the programme inherits the same control failures seen in other machine-identity environments. The challenge is not only what the agent finds, but how its own access is bounded and observed.
Practical implication: treat testing agents as non-human identities and apply explicit access scoping, logging, and revocation controls.
NHI Mgmt Group analysis
Hybrid pentesting is now a governance model, not a tooling preference. The central shift in this topic is that agentic automation and human assurance are solving different problems. Automation improves depth and continuity, while humans preserve defensibility, reviewer accountability, and audit acceptance. Organisations that treat this as a simple efficiency upgrade will miss the control design implications. The practical conclusion is that penetration testing governance must define where autonomous execution ends and human sign-off begins.
Agentic testing systems should be treated as governed non-human identities. Once a testing platform can choose actions and execute across a web attack surface, it is no longer just a tool, it is a delegated actor with authority boundaries. That creates an identity governance requirement around permissions, attribution, monitoring, and revocation. The article’s most important implication for IAM teams is that agent identity controls must extend beyond production workloads into security tooling itself.
Audit credibility depends on evidence quality, not just test volume. More tests are not automatically better if the output cannot be validated, signed, and understood by stakeholders. This is where security assurance programmes often fail: they optimise for speed or coverage and then discover that the artefacts are unusable for governance. The practical conclusion is that testing programmes need evidence standards as rigorous as their detection standards.
Assurance gap: The article highlights a common failure mode in modern security operations, where automated depth outpaces the organisation’s ability to validate and govern results. That gap becomes visible when testing output is technically rich but operationally weak, because no one can confidently attest to scope, accuracy, or safety. The practical conclusion is that governance should be designed around validated execution, not raw automation output.
The market is moving toward control-bound autonomy, not full autonomy. This topic signals that security leaders are unlikely to accept agentic systems that cannot show bounded execution and human oversight. For identity and security architects, that means the next buying and design questions will focus on delegation boundaries, review workflows, and auditable action trails. The practical conclusion is to evaluate agentic systems by how well they fit existing governance, not by how independently they operate.
What this signals
Agentic test agents create a governance signal as much as a security signal: once a tool can act independently, organisations need lifecycle discipline for its credentials, logging, and offboarding. The practical next step is to map those controls against established identity governance practices and the NIST Cybersecurity Framework 2.0.
The broader programme implication is that security testing is converging with identity governance. That means teams should review whether their non-human identity inventory, privileged access controls, and evidence workflows can already account for autonomous security tooling, or whether those controls only cover production workloads.
Assurance-bound autonomy: the emerging control pattern is not full autonomy, but autonomy that is bounded by reviewable evidence and human sign-off. Teams that can attribute actions, constrain scope, and revoke access quickly will be better positioned to operationalise agentic security without weakening audit confidence.
For practitioners
- Define execution boundaries for agentic test agents Limit target scope, allowed actions, and escalation paths before any agentic pentest run begins. Require explicit approval for changes to scope, and keep a revocation process ready if the agent exceeds its remit.
- Keep human validation in the reporting chain Make a reviewer responsible for confirming exploitability, safety, and report quality before results enter risk registers, board packs, or audit evidence. Human sign-off should be mandatory where outputs influence formal assurance.
- Treat testing platforms as non-human identities Assign unique credentials, least-privilege permissions, logging, and revocation to every autonomous testing workflow. Review these access paths the same way you review service accounts or other privileged machine identities.
- Separate detection depth from assurance evidence Measure coverage and exploit discovery separately from evidence quality, reviewer confidence, and audit usability. A programme is only working when both technical findings and governance artefacts are usable.
Key takeaways
- Agentic AI can expand pentest depth and continuity, but it does not remove the need for human-reviewed assurance.
- The main control issue is governance of the agent itself, including scope, attribution, and revocation.
- Security teams should judge these tools by whether they produce audit-ready evidence, not by how independently they operate.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
OWASP Agentic AI Top 10 and OWASP Non-Human Identity Top 10 address the attack and risk surface, while NIST CSF 2.0, NIST SP 800-53 Rev 5 and NIST AI RMF set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| OWASP Agentic AI Top 10 | Agentic testing systems need scope and oversight controls that map to agentic AI risk patterns. | |
| OWASP Non-Human Identity Top 10 | NHI-01 | The testing agent behaves like a governed non-human identity with delegated access. |
| NIST CSF 2.0 | PR.AC-4 | Access and permissions management is central to how testing agents are constrained. |
| NIST SP 800-53 Rev 5 | IA-5 | Authenticator management governs the credentials used by autonomous testing workflows. |
| NIST AI RMF | GOVERN | Governance of AI-enabled testing requires accountability, oversight, and clear decision rights. |
Apply agentic AI control reviews to constrain autonomous actions and require human approval for audit-facing output.
Key terms
- Agentic Pentesting: An approach to penetration testing that uses AI-driven systems to support planning, execution, or interpretation of tests. The key issue is not automation by itself, but whether the environment provides enough context for the output to be accurate, prioritised, and operationally useful.
- Audit-ready reporting: Patch reporting that is structured well enough to answer compliance questions without rebuilding data manually. It includes success, failure, exception, and affected-device evidence, and it should be exportable so auditors and internal control owners can verify that remediation really happened.
- Assurance gap: The mismatch between technical discovery and governance acceptance. A team can generate many findings and still fail to produce evidence that auditors, risk owners, or executives consider defensible, usually because validation, attribution, or review controls are incomplete.
- Assurance-bound autonomy: A control model where an autonomous system is allowed to act only inside explicit boundaries and with human review at defined checkpoints. It preserves the speed of automation while keeping accountability, evidence quality, and safety within governance limits.
What's in the full article
Terra's full article covers the operational detail this post intentionally leaves for the source:
- The exact hybrid pentesting workflow used to combine agentic exploration with human validation.
- The Terra Platform capabilities that support continuous testing, review, and audit-ready reporting.
- How the organisation balanced safety controls with deeper attack-surface coverage in practice.
- The deployment context on AWS and Amazon Bedrock that underpins the agentic testing workflow.
Deepen your knowledge
NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, machine identity security, and secrets management. It is designed for practitioners who need to apply identity control discipline across modern security and automation programmes.
Published by the NHIMG editorial team on August 2, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org