TL;DR: AI is pushing offensive security from point-in-time checks toward continuous, multi-stage testing, while FireCompass argues that most enterprises still assess only about 20% of their attack surface in annual pentests. The shift matters because attackers can now chain reconnaissance, credential discovery, and lateral movement faster than traditional testing cycles can validate exposure.
At a glance
What this is: This is FireCompass’s analysis of how AI and automation are changing penetration testing and red teaming, with the key finding that annual, partial-scope testing no longer matches the speed and breadth of AI-driven attack paths.
Why it matters: It matters to IAM and security teams because offensive testing increasingly needs to validate credential exposure, lateral movement, and privileged access pathways across human and non-human identities, not just static control checklists.
By the numbers:
- 90% of enterprises still conduct annual pentests covering only ~20% of the attack surface.
👉 Read FireCompass's keynote analysis of AI in offensive security and continuous red teaming
Context
AI-driven offensive security is best understood as a governance problem before it is a tooling problem. When attack planning, exploit generation, and path validation become faster and more adaptive, point-in-time testing stops reflecting the real exposure surface. The primary question for practitioners is no longer whether AI can assist red teams, but whether current assurance models can keep up with attack-path discovery across credentials, services, and privileged pathways.
The identity angle is real because modern offensive testing increasingly revolves around access. Credentials, service accounts, tokens, and elevated permissions are what make multi-stage attack chains possible, and that puts IAM, PAM, and NHI governance directly in scope. The article’s starting point is typical for organisations relying on periodic testing, but the scale of the gap is now more visible in AI-assisted attack environments.
Key questions
Q: How should security teams govern AI agents used for offensive testing?
A: Treat offensive AI agents as distinct workloads with explicit ownership, scoped tools, and logged approvals. Give them only the environments, credentials, and actions needed for authorised testing. Separate research targets from production systems, and review retries, data access, and output handling as part of standard governance, not as an afterthought.
Q: Why do annual penetration tests fall short against modern exploit timelines?
A: Annual testing assumes the attack surface stays stable long enough for point-in-time validation to remain representative. When vulnerabilities can be identified and weaponised far faster than that, gaps appear between assessments and the live environment. Continuous validation shortens that exposure window.
Q: What breaks when offensive testing does not include identity and privilege paths?
A: Teams can conclude that systems are safe while leaving credentials, service accounts, and elevated permissions untested. That creates blind spots where an attacker can pivot from a low-value exposure into a working compromise route. If identity paths are omitted, coverage becomes cosmetic rather than operational.
Q: How can organisations tell if offensive security is actually improving risk?
A: Look for shorter remediation cycles, fewer repeat findings, and better upstream decisions from engineering and security teams. If testing produces reports but does not change code quality, access patterns, or control design, it is generating evidence, not resilience.
Technical breakdown
Forward and backward chaining in AI attack planning
Forward chaining means the attacker starts from an observed state, such as an exposed service, credential, or open port, and expands step by step into a viable path. Backward chaining reverses the process, starting from the target, such as sensitive data or a privileged system, and then deriving the required prerequisites. In offensive security, these methods make attack planning more adaptive because the system can reason over multiple paths instead of following a fixed script. That is especially relevant in mixed cloud and identity environments where the shortest route is often through access, not software flaws.
Practical implication: test for attack paths, not just isolated vulnerabilities, especially where identity and privilege are the real control points.
What agentic AI changes in red team operations
Agentic AI is not just model output generation. It is a system that can select actions, decide timing, and execute tool use toward a goal, which makes it materially different from an LLM used for advice or content generation. In offensive workflows, that means an agent can enumerate assets, decide which path is promising, and continue without waiting for a human prompt at every step. The risk is not that the model is clever, but that the workflow becomes autonomous enough to scale repetitive attack validation across large environments.
Practical implication: treat agentic offensive workflows as active security systems that need scope, logging, and human oversight, not as passive analyst assistants.
Why continuous automated red teaming replaces annual coverage
Continuous automated red teaming is a coverage model, not just a faster pen test. It keeps validating the environment as assets, configurations, and identities change, which matters because attack surfaces now include ephemeral cloud services, shadow environments, and non-human identities that appear and disappear quickly. Annual testing cannot reliably capture that state drift. The deeper issue is assurance latency: if the exposure window changes weekly or daily, quarterly or annual assessment is already stale by the time results are reviewed.
Practical implication: align offensive testing cadence to environment change velocity, especially where access, secrets, and workloads are highly dynamic.
Threat narrative
Attacker objective: The attacker’s objective is to prove and exploit real attack paths from exposure to privilege, then reach sensitive systems or data before defenders notice the chain.
- Entry begins with externally reachable assets, exposed services, or credentials that an AI system can enumerate quickly across the attack surface.
- Escalation follows when the system maps those findings into multi-stage paths, using valid access or weak controls to move toward higher-value systems and permissions.
- Impact occurs when the path is validated end to end, showing that lateral movement and privilege misuse remain feasible despite periodic testing.
NHI Mgmt Group analysis
Continuous offensive validation is becoming an identity governance issue, not just a testing issue. Once attackers can chain discovery, credential access, and lateral movement at machine speed, the weak point is often not the vulnerability itself but the access model behind it. IAM, PAM, and NHI governance determine whether a discovered weakness becomes an actual breach path. Practitioners should treat offensive testing as an identity assurance loop, not a quarterly compliance exercise.
AI-driven attack planning exposes an attack-path visibility gap that most security programmes still underestimate. Traditional testing often measures whether a vulnerability exists, while AI-assisted adversaries care whether a path exists from exposure to impact. That distinction changes what “coverage” means in practice. Attack-path visibility: the ability to see how exposed assets, identities, and permissions connect into a working compromise route. Practitioners should prioritise path validation over isolated control reports.
Agentic AI in offensive security accelerates the convergence of tooling and decision-making. When a system can choose actions as well as execute them, the operational question shifts from “can it run a test” to “can it be trusted to stay inside scope.” That makes governance, logging, and containment part of the offensive stack. NIST AI RMF GOVERN and MANAGE functions are relevant here because the same oversight logic applies to AI systems that influence security decisions. Practitioners should define explicit runtime boundaries for any AI-assisted red team workflow.
The strongest signal in this article is the collapse of point-in-time assurance. Security teams still often assume that a successful review or annual assessment buys them a durable risk picture. AI-driven attack planning invalidates that assumption by shortening the time between exposure, path discovery, and exploitation. The practical conclusion is clear: controls that are not continuously validated are no longer trustworthy enough for dynamic environments.
What this signals
Assurance latency is now a programme-level risk. When exposures can be found and abused quickly, but remediation still takes weeks, offensive testing has to feed directly into control repair. The relevant question is not whether a vulnerability was seen, but how fast the organisation can revalidate identity-linked exposure after change. The gap between discovery and containment is where AI-assisted adversaries gain advantage.
Attack-path validation should become a standing part of IAM and PAM governance. Security teams that already manage privileged access, secrets, and workload identities need to connect those controls to red-team style validation. A control that is never exercised against real paths is only partially proven. For practitioners, that means aligning test cadence to credential turnover, infrastructure drift, and the emergence of shadow environments.
For practitioners
- Shift from point-in-time to continuous offensive validation Use automated red teaming or continuous attack simulation to re-test exposed assets, identity paths, and privilege chains whenever the environment changes, not just on a quarterly calendar.
- Prioritise identity-linked attack paths Map routes from exposed services to credentials, tokens, service accounts, and elevated permissions so testing focuses on the paths most likely to produce real compromise.
- Define scope boundaries for AI-assisted testing Constrain agentic testing systems with explicit asset scope, logging, approval rules, and stop conditions so autonomous actions stay inside authorised environments.
- Measure assurance latency, not just vulnerability counts Track how long it takes for a discovered exposure to be revalidated after a change, because stale findings are a stronger indicator of risk than raw scan volume.
Key takeaways
- AI-driven offensive security shifts assurance from periodic scans to continuous attack-path validation.
- The key failure is not only vulnerability presence, but untested identity and privilege routes from exposure to impact.
- Practitioners should bind autonomous testing to scope, logging, and remediation timing so coverage stays operational.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK and OWASP Agentic AI Top 10 address the attack and risk surface, while NIST AI RMF, NIST CSF 2.0 and NIST SP 800-53 Rev 5 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST AI RMF | GOVERN | The article centers on oversight for AI-assisted security workflows. |
| MITRE ATT&CK | TA0006 , Credential Access; TA0008 , Lateral Movement | The attack paths described rely on credential access and lateral movement. |
| NIST CSF 2.0 | PR.AC-4 | The article repeatedly returns to access and privilege as the control boundary. |
| NIST SP 800-53 Rev 5 | AC-6 | Least privilege is the core control tested by multi-stage attack simulation. |
| OWASP Agentic AI Top 10 | Agentic AI is part of the offensive workflow discussed in the article. |
Align offensive validation to PR.AC-4 by testing whether access entitlements actually block attack paths.
Key terms
- Agentic AI: Autonomous AI systems capable of planning, deciding, and taking actions — including calling APIs, writing code, and orchestrating other agents — with minimal human oversight. Agentic AI introduces new NHI risks as agents must authenticate to external services.
- Automated red-teaming: Automated red-teaming is the use of adversarial test generation to find how an AI model or agent fails under pressure. It goes beyond manual review by systematically probing prompt injection, goal drift, unsafe outputs, and other repeatable behavioural weaknesses before production use.
- Attack path: A sequence of identities, permissions, systems, and data stores that an attacker can traverse after obtaining trusted access. In practice, attack paths matter more than single accounts because they show how a low-risk identity can become a route to high-value exposure.
- Assurance Latency: Assurance latency is the delay between a control changing, a risk emerging, and the organisation recognising and acting on that change. Shorter latency means governance is closer to real conditions, which is critical for access, privilege, and machine identity controls.
What's in the full article
FireCompass's full article covers the operational detail this post intentionally leaves for the source:
- The keynote framing and examples behind continuous automated red teaming across full attack surfaces.
- FireCompass's explanation of forward and backward chaining in AI attack planning and how it changes red team workflow.
- The discussion of agentic AI versus LLMs in offensive security, including how autonomy changes execution.
- The practical examples of multi-stage attack simulation, including SMB enumeration, credential discovery, and lateral movement.
Deepen your knowledge
The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance and secrets management for practitioners who need to connect identity controls to real-world attack paths. It helps security teams strengthen the governance foundations that offensive testing keeps exposing.
Published by the NHIMG editorial team on September 3, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org