By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: TENZAIPublished July 29, 2026

TL;DR: Time-boxed testing now depends as much on machine-assisted discovery depth as on human exploitation skill, according to TENZAI. Markus Alliance says Tenzai helped it achieve 5x penetration testing efficiency, 4x higher attack surface coverage, and doubled engagement capacity by automating reconnaissance, discovery, mapping, and attack chaining across application tests.


At a glance

What this is: This is a case study on how hybrid AI and human penetration testing can expand discovery speed, coverage, and engagement capacity.

Why it matters: It matters to IAM and security practitioners because faster reconnaissance and broader attack surface mapping can reveal exposed credentials, privilege paths, and application access flaws earlier in the testing lifecycle.

By the numbers:

  • According to Markus Alliance, one engineer would typically reveal 20-25% of the attack surface in a week of discovery work for a large application.

👉 Read TENZAI's case study on hybrid AI-assisted penetration testing efficiency


Context

Application penetration testing is constrained by time, not just skill. When discovery and mapping consume most of an engagement, testers often reach the highest-risk areas too late, and low-priority paths remain unexamined. In this case study, the primary issue is not whether AI can replace testers, but whether AI can extend the portion of the application attack surface that humans can meaningfully validate within fixed timelines.

For IAM and identity-adjacent programmes, this has a direct control implication. Pen tests that surface user roles, API exposure, admin paths, and hidden attack chains can expose weaknesses in authorisation design, secrets handling, and privilege boundaries that traditional schedule-driven testing misses. The starting position described here is increasingly typical for teams operating under compressed delivery cycles and broad application portfolios.


Key questions

Q: How should security teams use AI-assisted penetration testing without losing trust in the results?

A: Use AI-assisted testing to widen discovery, then force a human validation step before any output becomes a confirmed finding. Teams should require traceable actions, repeatable evidence, and clear exploit paths so the machine is accelerating analysis rather than substituting for it. The output is most useful when it helps experts spend more time on high-impact validation.

Q: Why does broader attack surface coverage matter in application security programmes?

A: Broader coverage matters because time-boxed tests always leave some paths unexplored, and the missed paths are often where privilege boundaries, hidden APIs, and exposed secrets sit. If teams only measure speed or finding count, they can miss the control failures that let attackers move from discovery to access. Coverage is a governance signal, not just a test metric.

Q: What do security teams get wrong about AI-generated penetration testing findings?

A: The main mistake is treating AI output as proof rather than as a lead. Findings still need manual confirmation, especially when the issue involves chained weaknesses, session logic, or privilege escalation. Good programmes use AI to surface more candidate paths, then rely on experienced testers to prove whether those paths are real and material.

Q: How can organisations prioritise penetration testing when applications outnumber testers?

A: Focus on the applications and interfaces most likely to expose privileged access, sensitive data, or business-critical workflows, then use assisted discovery to broaden the first pass. That approach preserves depth where it matters while reducing time spent on repetitive reconnaissance. The key is to align test effort with risk, not with application count.


Technical breakdown

Autonomous discovery changes the front end of a penetration test

The front end of an application pen test is usually dominated by reconnaissance, endpoint mapping, role discovery, and basic traffic understanding. When an AI system performs those steps autonomously, it can assemble a first-pass model of the target in hours rather than days. That does not replace human judgment, but it changes the economics of test coverage by giving testers a structured view of endpoints, APIs, and user flows before manual exploitation begins. The practical effect is that the human team can spend more of the engagement on validation, chaining, and high-value hypothesis testing instead of repetitive discovery work.

Practical implication: test plans should assume machine-assisted discovery and reallocate human effort toward targeted exploitation and verification.

Attack chaining depends on the quality of intermediate findings

Attack chaining is the process of linking individually modest weaknesses into a higher-impact path, such as combining an exposed endpoint, weak role separation, and an access-control flaw. AI-assisted tools can accelerate the first pass across those relationships by identifying patterns faster than a manual review would. The risk is that teams may treat chain generation as proof rather than as a lead. In practice, the value comes from using AI to widen the candidate set, then using experienced testers to prove exploitability, blast radius, and business impact. That is especially relevant where identity, session state, and authorisation logic intersect.

Practical implication: validate chained findings with manual testing before treating them as confirmed risk paths.

Secrets exposure and privilege paths remain the most actionable application findings

The article points to timing attacks, source code analysis, and exposed secrets as examples of the weaknesses the platform helps surface. Those are not abstract findings. They often lead directly to credential abuse, admin-level access, or movement into underlying infrastructure. In identity terms, that means application security testing is increasingly a control for discovering weakly protected secrets, over-permissioned roles, and access logic that gives attackers a route from application exposure to privileged execution. The technical takeaway is that better discovery increases the odds of finding governance failures, not just code flaws.

Practical implication: treat discovered secrets and admin paths as identity-control failures, not only as application defects.


Threat narrative

Attacker objective: The objective is to convert broad application exposure into privileged access paths that demonstrate real business impact.

  1. Entry begins with autonomous reconnaissance against the target application, where the system maps endpoints, APIs, user roles, and exposed surfaces before human testers intervene.
  2. Escalation follows when discovered weaknesses are chained together, such as timing leaks, source code issues, or exposed secrets that can support higher-privilege access.
  3. Impact occurs when those paths produce admin-level access, broader infrastructure reach, or a validated attack chain that materially raises client risk.

NHI Mgmt Group analysis

Hybrid testing changes the unit of security work, but not the need for human judgment. AI-assisted reconnaissance can compress discovery from days into hours, yet the decisive security value still comes from skilled testers interpreting what the machine finds. That matters because automated breadth without human validation can produce noisy coverage rather than actionable risk reduction. Practitioners should treat AI as a discovery multiplier, not as an authority on exploitability.

Coverage compression is the named concept this case study illustrates: the gap between how much of an application can be mapped and how much can be meaningfully tested inside a fixed engagement window. Time-boxed assessments usually leave large parts of the attack surface unexplored, especially where discovery work consumes the schedule. When AI expands the first-pass map, it changes which weaknesses are even visible to the security team. The governance lesson is that coverage metrics now matter as much as finding counts; practitioners should measure whether their tests are reaching the right surfaces, not just completing the calendar.

Identity and authorisation failures are now first-class outputs of application testing. The most valuable findings in these workflows are not only code defects but exposed secrets, admin paths, weak role boundaries, and chained access failures. That is where appSec intersects directly with IAM and NHI governance. When testing surfaces credentials or privilege paths, the finding is as much about access control design as about software correctness. Practitioners should route these results into identity remediation, not just development backlogs.

The market is moving toward assisted security services rather than pure tool replacement. This case shows that customers value a blend of automation and expert review, especially when engagements need broader coverage without more headcount. That pattern is likely to reshape both testing services and internal assurance programmes. Teams should expect pressure to justify where human-only effort still adds value and where machine-assisted discovery is now table stakes.

Service design will increasingly be judged on evidence quality, not only speed. Faster assessments are only useful if they produce findings that withstand manual scrutiny and map to business risk. For security leaders, this creates a higher bar for procurement and for internal programmes: the question is not whether AI can do more work, but whether it helps teams reach defensible conclusions faster. Practitioners should demand transparent workflows and traceable findings.

What this signals

AI-assisted penetration testing will push more organisations to separate discovery from remediation more deliberately. If a machine can surface endpoints, roles, and secrets faster than a human can review them, the bottleneck moves to triage, proof, and ownership. That creates a governance problem as much as a tooling one, especially where application findings overlap with identity controls and secret handling. The operational question is no longer whether testing can be accelerated, but whether teams can absorb the findings without creating a larger backlog.

Coverage compression: the new testing constraint is no longer just time, but how much of the attack surface can be mapped before human validation begins. That will force practitioners to think harder about where assisted testing fits into assurance programmes and how much confidence they can place in partial coverage. For identity-adjacent findings, the signal should go into IAM, PAM, and secrets workflows rather than remaining isolated in an AppSec queue.


For practitioners

  • Re-baseline discovery coverage metrics Measure how much of the application attack surface is actually being mapped during the first phase of testing, including endpoints, APIs, user roles, and hidden workflows. Use that baseline to decide where AI-assisted discovery can extend coverage without lowering validation standards.
  • Separate AI-generated leads from confirmed findings Require a handoff step where human testers validate exploitability, business impact, and access scope before a finding is treated as actionable. Keep the machine output as evidence of discovery, not proof of compromise.
  • Route exposed secrets into identity remediation When testing reveals secrets, admin paths, or privilege shortcuts, send the issue to the teams that own authentication, authorisation, and secret handling. That prevents application findings from being closed as generic code defects when the real failure is access governance.
  • Use chained findings to prioritise high-value assets Rank chains that move from low-severity issues to admin-level access, infrastructure reach, or sensitive data paths above isolated low-risk findings. The goal is to focus remediation on the routes most likely to enable real attacker progress.

Key takeaways

  • AI-assisted testing increases discovery breadth, but human validation still determines whether a finding is actionable.
  • The most important outputs are often identity-adjacent, including exposed secrets, admin paths, and privilege chains.
  • Security leaders should judge testing programmes by coverage quality and remediation flow, not by speed alone.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

MITRE ATT&CK address the attack and risk surface, while NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
MITRE ATT&CKTA0007 , Discovery; TA0006 , Credential Access; TA0004 , Privilege EscalationThe article centers on autonomous discovery, exposed secrets, and privilege paths.
NIST CSF 2.0PR.AC-4The findings highlight access control and privilege boundary issues in applications.
NIST SP 800-53 Rev 5AC-6Least privilege is central when pen tests uncover admin-level access routes.
CIS Controls v8CIS-5 , Account ManagementAccount and role management weaknesses are part of the article's risk pattern.

Map automated discovery outputs to ATT&CK and prioritise validation of credential and privilege escalation paths.


Key terms

  • Attack Chain: An attack chain is a sequence of prompts, observations, and tool calls that moves an AI agent from a benign starting point to a harmful result. In agent security, the chain matters more than any single prompt because real risk often emerges only when actions accumulate across steps.
  • Attack Surface Coverage: Attack surface coverage is the share of a target system's reachable components that a test meaningfully examines. It is not just enumeration of assets. It reflects whether the testing process actually reaches the endpoints, workflows, and identities most likely to contain exploitable weakness.
  • Hybrid Penetration Testing: Hybrid penetration testing combines automated discovery or analysis with human exploitation and judgment. The machine expands breadth and speeds initial reconnaissance, while the tester validates impact, builds exploit chains, and decides which findings matter most for the client.
  • Privilege Boundary: A privilege boundary is the control line that separates ordinary user actions from elevated administrative actions. When the boundary is poorly enforced, attackers can repurpose normal tools or policy logic to cross into root-level execution without going through intended approval or validation steps.

What's in the full report

TENZAI's full case study covers the operational detail this post intentionally leaves for the source:

  • How the hybrid workflow is structured between autonomous discovery and human exploitation review
  • Examples of follow-on use with Ask Tenzai and Burp Suite extension workflows
  • The engagement model that supported a new low-touch penetration testing tier
  • Customer-facing transparency details showing how actions and findings were traced during testing

👉 The full TENZAI case study covers workflow transparency, follow-on testing, and service model changes.

Deepen your knowledge

The NHI Foundation Level course, the industry's only accredited NHI security programme, covers NHI governance, secrets management, and machine identity fundamentals. It helps practitioners align identity control decisions with the broader security programmes they run.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 1, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org