TL;DR: Project Glasswing, built around Anthropic’s unreleased Claude Mythos Preview model, found thousands of previously unknown zero-day vulnerabilities across major operating systems and browsers, including an OpenBSD bug reportedly missed for 27 years, according to Orca Security. Continuous AI-assisted testing is becoming practical, but human judgment still determines what gets fixed first.
Editorial analysis by NHI Mgmt Group, based on content published by Orca Security: “Anthropic’s Project Glasswing Is a Positive Step Toward Cleaner, Safer Production”.
Key questions
Q: How should security teams use AI-driven testing in the development lifecycle?
A: Security teams should place AI-driven testing inside normal development workflows so findings arrive before production, not after release.
Q: Why does continuous security testing reduce more than just vulnerability counts?
A: Because earlier discovery changes the quality of engineering decisions.
Q: What are the signs that AI-driven security testing is failing to stay safe and auditable?
A: Warning signs include agent actions that are hard to trace, testing that crosses approved scopes, and outputs that cannot be tied back to deterministic rules or role based permissions.
Practitioner guidance
- Embed security testing earlier in the SDLC Run AI-assisted investigation during design, implementation, and pre-release stages so developers can act before architectural choices harden into production risk.
- Use continuous testing for repeatable validation Replace reliance on one-off late-stage pen tests with recurring checks that developers and security teams can run throughout normal engineering workflows.
- Prioritise findings by reachability and impact Build triage criteria that weigh whether a weakness is actually exposed, reachable, and material to production risk before escalating it.
Bottom line: AI-driven security testing is moving vulnerability discovery closer to where software is built, which changes the economics of secure delivery.
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
AI-driven security testing is becoming a control-plane problem, not just an AppSec problem. The article frames the upside as better code review and earlier vulnerability discovery, but the deeper implication is that security testing is now part of the operational identity layer of software delivery. Once tools can inspect, triage, and trigger actions continuously, teams must decide which identities are allowed to initiate that work and under what authority. Practitioners should treat the testing stack as governed infrastructure, not a collection of helper scripts.
A few things that frame the scale:
- 91% of former employee tokens remain active after offboarding, leaving organisations vulnerable to potential security breaches, according to The 2025 State of NHIs and Secrets in Cybersecurity.
- 44% of NHI tokens are exposed in the wild, being sent or stored over platforms like Teams, Jira tickets, Confluence pages, and code commits, according to The 2025 State of NHIs and Secrets in Cybersecurity.
A question worth separating out:
Q: How do organisations know if continuous security testing is actually working?
A: It is working when more issues are found before release, fewer emergency fixes reach production, and engineering spends less time on avoidable investigations. The best signal is not the number of findings alone, but whether the programme is reducing downstream noise and shortening remediation paths.
👉 Read our full editorial: AI-driven security testing is reshaping secure software delivery
AI-driven testing is changing the economics of pre-production vulnerability discovery. When security investigation becomes more continuous and easier to run, the old assumption that meaningful testing must be late and scarce starts to break down. That does not eliminate human review, but it does move more of the security decision surface into the development cycle. The practitioner conclusion is straightforward: teams should treat security testing as an operating rhythm, not an end-of-pipeline checkpoint.
A question worth separating out:
Q: Should organisations still keep pen testing if AI can test code continuously?
A: Yes. Continuous AI-assisted testing expands coverage, but it does not replace human judgment or deeper contextual review. Pen testing still matters where teams need adversarial thinking, business context, and confirmation of real-world exposure rather than only static defect discovery.
👉 Read our full editorial: AI-driven security testing is reshaping secure software delivery