By NHI Mgmt Group Editorial TeamBased on Orca Security: “Anthropic’s Project Glasswing Is a Positive Step Toward Cleaner, Safer Production” (April 13, 2026)

TL;DR: Project Glasswing, built around Anthropic’s unreleased Claude Mythos Preview model, found thousands of previously unknown zero-day vulnerabilities across major operating systems and browsers, including an OpenBSD bug reportedly missed for 27 years, according to Orca Security. Continuous AI-assisted testing is becoming practical, but human judgment still determines what gets fixed first.


At a glance

What this is: This analysis argues that AI-driven security testing is making earlier and more continuous vulnerability discovery practical inside the software development lifecycle.

Why it matters: For AppSec and engineering teams, the shift matters because it can reduce late-stage noise, improve deployment confidence, and move security validation closer to where code is built.


Context

AI-driven security testing is changing where vulnerability discovery happens in the software development lifecycle. Instead of relying mainly on late-stage testing, teams can now use automated investigation earlier in design, implementation, and release preparation.

The governance gap is not whether testing exists, but whether it is continuous enough to influence engineering decisions before code reaches production. That matters for application security, delivery speed, and how teams prioritise remediation when findings arrive.

This article uses Anthropic’s Project Glasswing as the trigger for a broader question: what happens when security testing becomes cheaper, more frequent, and more embedded in development workflows? The answer is less about replacing human reviewers and more about shifting where judgment is applied.


Key questions

Q: How should security teams use AI-driven testing in the development lifecycle?

A: Security teams should place AI-driven testing inside normal development workflows so findings arrive before production, not after release. The useful pattern is continuous validation during design, build, and pre-release review, paired with human judgment for prioritisation. That reduces late-cycle noise and makes remediation cheaper because the code context is still fresh.

Q: Why does continuous security testing reduce more than just vulnerability counts?

A: Because earlier discovery changes the quality of engineering decisions. It lowers the number of issues that reach production, reduces urgent downstream escalations, and gives security teams a cleaner signal to triage. The value is not volume alone, but fewer exploitable problems escaping into live systems.

Q: What are the signs that AI-driven security testing is failing to stay safe and auditable?

A: Warning signs include agent actions that are hard to trace, testing that crosses approved scopes, and outputs that cannot be tied back to deterministic rules or role based permissions. If teams cannot tell what the agent did, why it did it, and under whose authority it operated, the control model is too loose. Safe testing should leave a clear audit trail from task assignment through execution.

Q: Should organisations still keep pen testing if AI can test code continuously?

A: Yes. Continuous AI-assisted testing expands coverage, but it does not replace human judgment or deeper contextual review. Pen testing still matters where teams need adversarial thinking, business context, and confirmation of real-world exposure rather than only static defect discovery.


Technical breakdown

How AI-assisted testing changes vulnerability discovery

AI-assisted security testing uses models to inspect code, surface suspicious logic, and explore likely exploit paths at a pace that manual testing cannot match. In practice, this turns vulnerability discovery into a more continuous activity instead of a late-stage event. The technical value is not just speed. It is the ability to run more checks across more code paths, earlier in the lifecycle, before design assumptions harden into production risk.

Practical implication: treat AI-assisted testing as an upstream discovery layer, not a replacement for human review.

Why late-stage pen testing leaves gaps

Traditional penetration tests and red-team exercises remain valuable, but they are still point-in-time activities. They often arrive after architecture, implementation, and release decisions are already locked in, which means the cost of fixing findings is higher and the feedback loop is slower. This creates a structural gap between finding a weakness and influencing the code path that created it. Continuous testing narrows that gap by making security feedback available while the software is still moving.

Practical implication: move testing left enough that findings can still change implementation choices.

Why context and prioritisation still decide outcomes

Finding a vulnerability is not the same as deciding what matters. AI can increase the volume of findings, but security teams still need context about reachability, business impact, and operational exposure. That is why judgment remains central even when detection becomes cheaper. The best technical outcome is a cleaner signal, not just a larger queue. In mature programmes, investigation data should feed prioritisation, not overwhelm it.

Practical implication: pair automated discovery with triage rules that rank exposure, not just defect count.


NHI Mgmt Group analysis

AI-driven testing is changing the economics of pre-production vulnerability discovery. When security investigation becomes more continuous and easier to run, the old assumption that meaningful testing must be late and scarce starts to break down. That does not eliminate human review, but it does move more of the security decision surface into the development cycle. The practitioner conclusion is straightforward: teams should treat security testing as an operating rhythm, not an end-of-pipeline checkpoint.

The strongest value is not just AppSec coverage, but better security signal quality. Catching defects earlier reduces production noise, urgent escalations, and the downstream cost of triage. That matters because many organisations still measure success by the number of findings closed rather than by how early they influence engineering choices. The real governance win is fewer exploitable issues entering production in the first place. Practitioners should judge these tools by whether they improve decision quality, not by whether they generate more alerts.

Continuous testing narrows the gap between code risk and operational risk, but it does not remove the need for context. Static findings only become useful when teams know what is reachable, what is exposed, and what would actually matter in production. That is why secure software delivery now depends on combining automated investigation with human prioritisation. The implication for programmes is clear: invest in tooling that improves signal, then make sure your triage model can translate findings into action.

Continuous security validation is becoming a delivery capability, not a specialist activity. The market signal here is that development teams increasingly need security checks that fit normal workflows rather than special projects. That does not mean every engineer becomes a security analyst. It means the process of finding and fixing issues has to be lightweight enough to happen repeatedly without slowing delivery to a crawl. Practitioners should expect security validation to move closer to build and release automation.

Judgment remains the differentiator as automation expands. The easy part is surfacing possible issues. The hard part is deciding what to fix first, what can wait, and what signals real exposure. As more organisations adopt AI-assisted testing, the security function shifts from discovery alone toward prioritisation and contextual decision-making. The programme implication is that teams will need stronger review criteria, not just more findings.

What this signals

AI-assisted testing is becoming a delivery control, not just an AppSec enhancement. The practical shift is that vulnerability discovery can now happen often enough to influence how code is written, reviewed, and released. That changes the operating model for engineering and security teams: the question becomes whether the programme can absorb frequent findings without turning triage into a bottleneck.

Earlier testing only pays off when prioritisation remains disciplined. If automation increases the number of findings faster than teams can assess reachability and business impact, the result is more noise, not better security. NHI Mgmt Group’s view is that mature programmes will measure success by the quality of decisions improved, not by the raw number of issues uncovered.


For practitioners

  • Embed security testing earlier in the SDLC Run AI-assisted investigation during design, implementation, and pre-release stages so developers can act before architectural choices harden into production risk.
  • Use continuous testing for repeatable validation Replace reliance on one-off late-stage pen tests with recurring checks that developers and security teams can run throughout normal engineering workflows.
  • Prioritise findings by reachability and impact Build triage criteria that weigh whether a weakness is actually exposed, reachable, and material to production risk before escalating it.
  • Keep human review in the decision loop Use automation to expand coverage, then require human judgment to decide what gets fixed first and what evidence is sufficient to move an issue forward.
  • Measure security by production quality Track whether earlier testing reduces downstream noise, late escalations, and the volume of issues that escape into live environments.

Key takeaways

  • AI-driven security testing is moving vulnerability discovery closer to where software is built, which changes the economics of secure delivery.
  • The main benefit is not only more findings, but fewer late-stage surprises and less downstream noise in production.
  • Teams still need human judgment to rank what matters, because automated discovery does not by itself explain impact or urgency.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP SAMM, OWASP ASVS, NIST CSF 2.0 and CIS Controls v8 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP SAMMSecurity TestingThe article is about shifting security testing earlier and making it continuous in the SDLC.
Recommendation — Embed recurring security testing into the SDLC and use findings to change engineering decisions earlier.
OWASP ASVSV15 — Secure Coding and ArchitectureContinuous testing is aimed at finding weaknesses before they become architectural defects in production.
Recommendation — Use ASVS V15 to verify secure design choices before code is promoted.
NIST CSF 2.0PR.DS-10 — Data-in-Transit Is ProtectedThe article concerns secure software delivery, where validation supports protective controls across the build-to-prod path.
Recommendation — Map upstream testing improvements to protective controls that reduce live-environment exposure.
CIS Controls v8CIS-16 — Application Software SecurityThe piece centres on application security testing as part of software delivery.
Recommendation — Apply CIS-16 to formalise security testing across the development lifecycle.

Key terms

  • Shift-left testing: Security testing that happens earlier in the development lifecycle, usually before production deployment. It reduces exposure windows and creates earlier evidence of control execution, which is especially valuable when the organisation needs to show systematic risk management.
  • Continuous Security Testing: A security model that revalidates an AI agent whenever its prompt, model, tools, memory, or permissions change. For agentic systems, this is not a pipeline stage but a living control that tracks behaviour as the system evolves in production.
  • Security Signal: An observable event or alert produced by a security control that may indicate malicious activity, misconfiguration, or policy violation. Security signals are only useful when they can be collected, correlated, and acted on quickly. In low-telemetry or offline environments, signal quality often determines whether detection is practical at all.
  • Pre-production Vulnerability Testing: Pre-production vulnerability testing is security validation performed before code reaches live users. In PCI contexts, it is used to prove that exploitable flaws are detected early enough to prevent release, reduce remediation cost, and create audit evidence that testing happened continuously rather than occasionally.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 9, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org