By NHI Mgmt Group Editorial TeamBased on Aqua Security: “When AI Writes, Scans and Fixes Code, Runtime Becomes the Last Line of Defense” (February 26, 2026)

TL;DR: AI-driven code security can now find more than 500 high-severity vulnerabilities in production open-source codebases, according to Aqua Security citing Anthropic’s Claude Code Security results. That improves pre-deployment analysis, but it does not replace runtime security, because production-only misconfigurations, privilege escalation, and workload drift still determine real exposure.


At a glance

What this is: Aqua Security says AI code scanning can expose severe flaws in source, but it still cannot see the production behaviour that determines real exposure.

Why it matters: This matters because IAM and security teams still need runtime controls for workloads, privileges, and deployment drift even when AI improves pre-production code review.


Context

AI code scanning is the use of models to reason about code paths, data flow, and flaw patterns before software reaches production. The governance gap is that better pre-deployment analysis does not change what happens once workloads are assembled, configured, and granted runtime privileges.

For identity and cloud security teams, the real question is not whether code review gets smarter. It is whether runtime security still has to carry the controls that source analysis cannot see, including container behaviour, privilege escalation, and environment drift.

Aqua Security frames that boundary as architectural rather than temporary. That framing is accurate for most modern application stacks, where the operating context, not the source tree alone, determines exposure.


Key questions

Q: Why do AI code scanners still leave runtime security gaps?

A: Because scanners can reason over source, history, and known flaw patterns, but they cannot run the application in its production context. They miss whether a vulnerability is reachable through live authentication chains, whether a container spawns dangerous processes, and whether deployment settings make the issue exploitable.

Q: How should teams decide which code findings matter most in production?

A: Prioritise findings by runtime reachability, active privileges, exposed network paths, and whether the vulnerable component is actually loaded in the live workload. A finding that is severe in source may be low priority if production context makes it unreachable, while a smaller flaw can become urgent if it is directly exposed.

Q: What breaks when organisations rely on shift-left alone?

A: They lose sight of the layer where software is assembled, configured, and granted real permissions. That leaves misconfigurations, privilege escalation paths, dependency drift, and container behaviour outside the control model, even if pre-deployment scanning is highly effective.

Q: What should security teams do when code velocity increases faster than review capacity?

A: Move enforcement and triage closer to production so the security programme can see what actually executes, connects, and changes. Runtime telemetry becomes the deciding evidence for whether a new deployment expands exposure or simply adds noise.


Technical breakdown

Why source code analysis stops at the deployment boundary

Static and reasoning-based scanners can trace source, commit history, and known flaw patterns, but they cannot execute requests through a live API stack or observe a running container. That means they miss the behaviour that only appears when authentication middleware chains together, when processes spawn under real load, or when a workload touches files and network paths at runtime. The important point is not that scanners are weak. It is that they operate before the application is assembled into its real production state.

Practical implication: Treat pre-deployment scanning as a code-risk control, not as evidence that runtime exposure has been eliminated.

Why runtime behaviour changes the meaning of a vulnerability

A vulnerability in source code is not the same as a reachable exposure in production. Runtime context shows whether a package is loaded, whether a code path is actually executed, whether a privilege escalation path exists, and whether a misconfiguration makes the flaw exploitable. That distinction matters because many incidents come from the surrounding stack: base images, libraries, deployment settings, and workload permissions, not only the application logic itself.

Practical implication: Use live workload context to decide which findings are exploitable now rather than treating every code finding as equal risk.

Why AI amplifies the shift-left paradox

AI improves code discovery and remediation, but it also increases the volume and speed of change. More code reaches production faster, more dependencies move, and more configuration shifts occur between build and runtime. That compresses the window in which pre-deployment controls can tell the full story. The result is a sharper split between what can be proven in code and what can only be verified in production telemetry.

Practical implication: Plan for faster production drift and design runtime controls to absorb the exposure that pre-deployment tooling cannot model.


Read and download The State of NHI & AI Agent Breach Report 2026, covering 200+ breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Runtime context is now the governing truth for modern application risk: AI-assisted code analysis can improve flaw discovery, but it does not establish whether a vulnerability is reachable in a live workload. The decisive question moves from "what is in the code" to "what is actually executing, connected, and privileged in production." Practitioners should treat runtime observability as the control plane for exposure decisions.

Pre-deployment automation narrows the visible risk set, but it also concentrates attention on the residual runtime gap: once code scanning gets better, the remaining issues are more likely to be environment-specific misconfigurations, workload drift, and privilege paths that source analysis cannot infer. That makes runtime protection the place where security teams distinguish theoretical findings from actionable exposure.

Identity blast radius is the named concept this debate exposes: AI can accelerate code security, but it cannot tell you how far a compromised or misconfigured workload can move once identities, tokens, and network paths are active in production. The practical implication is that exposure is defined less by code quality than by the reach of runtime privilege.

Shift-left does not disappear, it becomes incomplete without shift-right: the industry has spent years pushing detection earlier, yet production remains where dependencies, configurations, and execution paths become real attack surfaces. This is the point at which NIST-CSF protection and detection outcomes, rather than code review alone, determine whether the programme is actually reducing risk.

Runtime security is the control that preserves accountability when AI speeds delivery: if code changes arrive faster and with less human inspection, then governance has to move to live behaviour, not just artefact inspection. Teams should expect the centre of gravity in cloud native security to continue moving toward production telemetry and enforcement.

From our research library:

  • Companies are dedicating an average of 32.4% of their security budgets to secrets management and code security, with US organisations leading at 40.8%, according to the State of Secrets in AppSec.
  • Claude Code-assisted commits leaked secrets at a rate of 3.2%, more than double the human-only baseline of 1.5%, with peaks reaching 31 secrets per 1,000 commits in August 2025, according to the State of Secrets Sprawl 2026.
  • Read next: NHI Lifecycle Management Guide

What this signals

Identity blast radius: once AI speeds code delivery, the governance problem is no longer just code quality but how far a workload can move when credentials, tokens, and network reach are live in production. Security teams need to evaluate exposure at issuance and runtime, not only at scan time.

Runtime analysis changes prioritisation by showing which findings are actually reachable in the current environment. That is the point where vulnerability management, workload identity, and cloud runtime controls converge into one operational decision set.


For practitioners

  • Define the runtime security boundary Separate issues that can be resolved in code review from issues that only become visible once the workload is running, connected, and privileged.
  • Prioritise exploitable findings over theoretical ones Use live workload context to rank vulnerabilities by reachability, active processes, network exposure, and deployment state rather than raw scanner output.
  • Inspect workload drift and privilege changes Track when containers spawn unexpected processes, touch new files, or gain permissions that were not present in the approved image.
  • Treat third-party dependencies as runtime risk Assume upstream packages and base images can create exposure windows that pre-deployment scanning cannot close, especially before patches are available.
  • Build enforcement around production telemetry Feed runtime signals into detection, triage, and remediation so security decisions are based on what is actually happening in the environment.

Key takeaways

  • AI improves code-level detection, but it does not resolve the production-side conditions that determine whether a flaw is reachable.
  • The main evidence in this article is the boundary between source analysis and runtime behaviour, where misconfiguration and privilege escalation still occur.
  • Security teams should keep runtime controls central because AI-driven development makes production telemetry more, not less, important.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 provides the primary governance reference for this term.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-05 — Overprivileged NHILive workloads can become overprivileged even when source code looks clean.
Recommendation — Review workload permissions against NHI-05 and reduce live privilege scope where production telemetry shows unnecessary reach.

Key terms

  • Runtime Security: Runtime security is the practice of detecting and constraining malicious behavior while software is executing. It focuses on live workload activity, not just code quality or pre-deployment checks, so teams can contain abuse after a system is already running.
  • Shift Left: Shift left is the practice of moving security checks earlier in the software development lifecycle, usually into planning, code review, or build steps. It reduces some defects before deployment, but by itself it does not control runtime behaviour, identity sprawl, or secrets misuse after release.
  • Workload Drift: Workload drift is the gap between an approved image or configuration and the software that actually runs in production. It can include changed packages, unexpected processes, altered network behaviour, or permissions that appear only after deployment.
  • Reachability analysis: Reachability analysis checks whether a vulnerability can actually be exploited in the application’s real code paths and dependency graph. It helps teams distinguish theoretical findings from issues that an attacker can reach, which makes prioritisation far more accurate for both AppSec and identity risk management.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 24, 2026.
Updated on October 11, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org