TL;DR: Continuous, autonomous AI pentesting is moving from conference buzz to a real buying category, as 34% of Latio respondents named “AI Pentesting” as the AI feature they are most excited about and Gartner cited the space in its continuous offensive security testing guidance, according to Novee. The real issue is no longer whether to adopt it, but whether a platform can replicate attacker behaviour against LLMs, copilots, and agents rather than simply automate familiar scanning workflows.
NHIMG editorial — based on content published by Novee: RSAC 2026 offers a glimpse into the future of offensive security
Questions worth separating out
Q: What breaks when AI pentesting only automates scanner workflows?
A: Teams get output that looks like offensive testing but does not prove attacker behaviour.
Q: Why do AI agents complicate existing IAM and NHI controls?
A: They complicate control design because they can select actions at runtime, call multiple APIs, and move authority across systems without a human session boundary.
Q: How do security teams know whether an AI pentesting tool is credible?
A: Ask whether it can show multi-step attack chains that begin with an actual entry condition and end with a validated impact.
Practitioner guidance
- Classify AI systems by access behaviour Inventory chatbots, copilots, and agents by whether they only generate text or can also call tools, read data, and invoke workflows.
- Demand adversarial evidence, not feature claims Evaluate AI pentesting platforms on whether they can demonstrate prompt injection chains, tool abuse, and validated exploit paths against realistic targets.
- Separate model risk from access risk in assessments Run AI security reviews in two layers: one for model behaviour such as jailbreaks and one for delegated access such as secrets, API permissions, and downstream tool action.
What's in the full article
Novee's full article covers the operational detail this post intentionally leaves for the source:
- The live RSAC context around how buyers are separating autonomous testing from legacy scanning labels.
- The specific demo flow for mapping an attack chain from a domain name to validated findings and remediation output.
- The buyer questions the vendor says matter most when comparing AI pentesting platforms.
- The conference sessions cited as market signals for where offensive AI security is heading.
👉 Read Novee's RSAC 2026 breakdown of AI pentesting and offensive security →
AI pentesting is maturing fast. What should teams evaluate now?
Explore further
Continuous AI pentesting is becoming a governance problem, not just a testing category. Once AI systems can reason, retrieve, and act through tools, the question shifts from finding bugs to proving whether an attacker can steer an operational workflow. That creates overlap between application security, IAM, and NHI governance, especially where service credentials or delegated tokens are involved. Practitioners should treat testing evidence as a control signal, not a marketing claim.
A question worth separating out:
Q: How should organisations govern AI agents alongside human identity and device access?
A: Organisations should treat AI agents as a separate identity class with their own entitlement boundaries, logging expectations, and approval model. Human IAM controls often assume interactive sign-in and review cycles, which do not fit autonomous or programmatic access. The safer approach is to define actor-specific policy and verify which access paths can be delegated without expanding trust unnecessarily.
👉 Read our full editorial: Continuous AI pentesting is redefining offensive security testing