By NHI Mgmt Group Editorial TeamDomain: Cyber SecuritySource: XbowPublished July 22, 2026

TL;DR: When AI models can surface vulnerabilities at scale, raw finding counts stop distinguishing noise from exploitable risk, according to Xbow's whitepaper. Security teams need metrics that reflect attackability, remediation progress, and operational priority, not just scan coverage or critical issue totals.


At a glance

What this is: This whitepaper argues that AI-assisted vulnerability discovery makes traditional AppSec reporting less useful because finding volume no longer tracks meaningful risk.

Why it matters: It matters to IAM and security practitioners because the same metric problem appears whenever automation expands discovery faster than governance can triage, prioritise, and assign ownership.

👉 Read Xbow's whitepaper on how AI is changing AppSec metrics


Context

Application security metrics break down when discovery scales faster than remediation. In an AI-driven environment, teams can generate more alerts, more findings, and more false urgency, but those counts alone say little about exploitability, exposure, or whether a control is actually reducing risk.

That matters beyond AppSec. The same governance problem shows up in IAM, NHI, and agentic AI programmes when teams measure volume instead of control outcomes. If a metric does not help identify what can be abused, what can be reached, or what can be remediated first, it is mostly reporting noise.


Key questions

Q: How should security teams measure AppSec effectiveness when AI tools surface far more findings?

A: Measure whether the programme is reducing exploitable exposure, not just producing more detections. Useful signals include time to remediate reachable issues, the percentage of critical findings with verified exploit paths, and the rate at which high-risk flaws are closed before release. Raw finding counts should support operations, but they should not define success.

Q: Why do traditional AppSec metrics become less useful when AI improves vulnerability discovery?

A: Because discovery volume can rise faster than remediation capacity, which makes counts and scan percentages drift away from actual risk. A stronger model separates what was found from what can be abused and from what was fixed. That gives teams a clearer view of attackability, ownership, and whether security work is shrinking real exposure.

Q: What do security teams get wrong about appsec metrics?

A: They often measure the number of vulnerabilities found instead of the speed and consistency of remediation. High finding counts can reflect better detection, not better security. The more useful signals are time to remediate, closure rates, and whether teams are fixing issues in the workflow that produced them.

Q: How do teams know whether AI-assisted AppSec is actually helping?

A: Look for findings that can be traced back to named components, repeated across assessments, and mapped to concrete remediation actions. If the system produces faster output but reviewers still cannot understand why a requirement exists, the programme has improved throughput without improving governance.


Technical breakdown

Why finding counts stop working as a security signal

Finding volume is a discovery metric, not a risk metric. AI tools can increase the number of issues surfaced without changing the underlying attack surface in a proportional way. That creates a measurement gap: teams see more output, but not necessarily more actionable risk. In practice, the same codebase can look worse simply because the scanner is better at detecting weak points. The right response is to separate discovery from prioritisation and track whether a finding is reachable, exploitable, and tied to an asset or workflow that matters.

Practical implication: move from raw issue totals to risk-ranked queues that distinguish discovered defects from exploitable exposure.

How AI changes vulnerability discovery for attackers and defenders

AI compresses the time and effort needed to identify weaknesses, which changes both sides of the equation. Defenders can discover more, but attackers can also identify promising targets faster and at greater scale. That means security programmes need metrics that reflect adversarial efficiency, not just tool output. A useful AppSec programme should tell you whether remediation is keeping pace with discovery, whether high-risk issues are closing faster than new ones appear, and whether teams are reducing the pool of materially exploitable flaws rather than just shrinking a backlog.

Practical implication: measure attacker-relevant exposure reduction, not only scanner throughput or backlog size.

What exploitable-risk metrics should replace vanity coverage metrics

Better metrics focus on decision quality. Examples include time to remediate exploitable findings, percentage of critical issues with verified exploit paths, and the share of high-risk findings that remain open across release cycles. These indicators are more useful than code-scan percentages because they show whether the programme is reducing meaningful exposure. In identity-heavy environments, the same logic applies to secrets, service accounts, and access pathways: count what can actually be abused, not just what was detected. Mature programmes link findings to ownership, business impact, and closure evidence.

Practical implication: align reporting to exploitability, remediation speed, and closure evidence rather than scan completeness.


NHI Mgmt Group analysis

Finding abundance creates a governance illusion if teams keep measuring output instead of exposure. AI-driven AppSec tooling can produce more detections, but more detections do not automatically mean less risk. The programme may look busier while the most exploitable issues remain open. Practitioners should treat raw finding counts as a hygiene metric, not an outcome metric.

Exploitable-risk scoring is the named concept this market needs more clearly. The article points to a shift from counting defects to weighting them by reachability, blast radius, and closure status. That is the only way to distinguish noise from security improvement when discovery scales faster than remediation. Teams should build reporting around exploitability, not volume.

Identity-adjacent security programmes should recognise the same failure pattern in secrets, service accounts, and access paths. Once AI accelerates discovery, the important question becomes which exposed credential, token, or access path can actually be abused. That is directly relevant to NHI governance because secrets and machine identities often create the shortest path from discovery to compromise. Practitioners should connect AppSec metrics to identity-owned risk where those pathways intersect.

Metric reform matters because board reporting needs fewer counts and more decision-grade signals. Leaders do not need another dashboard that tallies issues. They need metrics that show whether remediation is reducing attackability, whether high-risk exposures are closing faster than new ones appear, and whether ownership is clear enough to keep pace. Teams should make the metric set defensible to both engineering and governance audiences.

What this signals

Exploitability-weighted reporting will become the more durable operating model for AppSec, and the same logic will spread into NHI and secrets governance where exposure windows matter more than issue counts. Teams that still present scan totals as proof of control maturity will struggle to explain risk to both engineering and leadership.

As AI increases discovery speed, security programmes need a second layer of telemetry that answers whether exposure is shrinking or simply being counted more efficiently. That will push practitioners toward remediation SLAs, ownership clarity, and control effectiveness measures that can survive board scrutiny.

The practical test is whether metrics help teams choose actions faster. If a dashboard cannot distinguish a reachable secret from a theoretical defect, or a closed issue from a reopened one, it is not a governance instrument.


For practitioners

  • Replace raw finding counts with exploitability-weighted reporting Track whether each issue is reachable, weaponisable, and tied to an asset or workflow that matters. Use that view for executive reporting instead of scan volume or critical issue totals.
  • Measure remediation by closure evidence, not backlog size Report time to remediate exploitable findings, the percentage closed with verified fixes, and the number of high-risk issues carried across release cycles.
  • Tie AppSec metrics to identity-owned exposure Include secrets, service accounts, API keys, and other machine access paths in the same prioritisation model so AppSec does not ignore the shortest route to compromise.
  • Create a separate signal for discovery and prioritisation Keep scan coverage, issue volume, and triage queues distinct so a better detection engine does not masquerade as a weaker security posture.

Key takeaways

  • AI-assisted discovery makes raw AppSec finding counts less useful because volume no longer equals risk.
  • The most useful metrics now weight findings by exploitability, remediation speed, and verified closure.
  • Identity-adjacent exposures such as secrets and service accounts should be included in the same prioritisation model as application flaws.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, NIST SP 800-53 Rev 5, CIS Controls v8 and NIST AI RMF set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0ID.RA-1Risk assessment and prioritisation are central to the metric shift discussed here.
NIST SP 800-53 Rev 5SI-2Flaw remediation and timely correction align with the article's focus on meaningful closure.
CIS Controls v8CIS-7 , Continuous Vulnerability ManagementContinuous vulnerability handling depends on prioritisation, not just discovery.
NIST AI RMFMEASUREAI-driven discovery changes how teams measure programme effectiveness and risk reduction.

Use MEASURE to define metrics that show whether AI-driven security work reduces attackability and improves decision quality.


Key terms

  • Exploitable risk: Exploitable risk is the subset of discovered issues that an attacker can realistically use to gain access, move laterally, or cause impact. It depends on reachability, privilege, exposure, and business context, not on severity labels alone.
  • Finding volume: Finding volume is the total number of issues surfaced by scanners, tests, or AI-assisted discovery tools. It is a useful operational signal, but by itself it says little about whether the programme is actually reducing the organisation's attack surface.
  • Closure evidence: Closure evidence is proof that a vulnerability or control gap has been genuinely remediated and not just marked complete. It can include fixed code, validated config changes, retesting results, or control assertions that show the risk is no longer active.
  • Exploitability-weighted reporting: Exploitability-weighted reporting is a measurement model that ranks findings by how likely they are to be used in an attack and how much harm they could cause. It helps security teams avoid treating all discoveries as equally urgent.

What's in the full report

Xbow's full whitepaper covers the operational detail this post intentionally leaves for the source:

  • Specific metric examples for comparing exploitability, remediation progress, and risk reduction in AppSec programmes
  • Guidance on how AI changes attacker and defender workflows without relying on simple finding counts
  • A practical framework for evaluating whether security metrics are measuring exposure or just scanner output
  • Implementation-oriented advice for teams rethinking AppSec dashboards and executive reporting

👉 Xbow's full whitepaper covers the metric framework, prioritisation logic, and reporting changes in more detail

Deepen your knowledge

NHI Mgmt Group covers identity security, NHI governance, and agentic AI through independent research, practitioner guides, and the NHI Foundation Level course, the industry's only accredited NHI security programme. It is designed for practitioners who need a governance lens across human and machine access.
NHIMG Editorial Note
Published by the NHIMG editorial team on August 11, 2026.
NHI Mgmt Group — the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org