TL;DR: AI pentesting can expand testing speed and coverage, but Synack argues that the real challenge is not the demo result, it is the reliability engineering, human validation, infrastructure and model dependency required to run it at scale. The economics of repeated AI workflows, not just initial development, determine whether build-vs-buy holds up in production.
NHIMG editorial — based on content published by Synack: AI Pentesting Works. Building It Yourself Is the Hard Part
By the numbers:
- 91% of former employee tokens remain active after offboarding, leaving organisations vulnerable to potential security breaches.
- 44% of NHI tokens are exposed in the wild, being sent or stored over platforms like Teams, Jira tickets, Confluence pages, and code commits.
Questions worth separating out
Q: How should security teams decide whether to build or buy AI pentesting capabilities?
A: Teams should compare the full operating cost, not just the first prototype.
Q: Why do AI agent workflows need identity governance for oversight?
A: Because oversight only works when the organisation can prove who approved an action, what they saw, and why they intervened.
Q: What do security teams get wrong about AI safety testing?
A: The common mistake is treating AI safety testing as if it were just another security scan.
Practitioner guidance
- Build a full run-cost model Estimate tokens, compute, orchestration, monitoring, validation and human review for repeated AI pentesting runs rather than a single pilot.
- Treat foundation models as governed dependencies Document model providers, API dependencies, version changes and retirement triggers in the same inventory used for critical services.
- Separate generation from validation Use AI to accelerate coverage, but keep an independent validation layer for exploitability, reproducibility and business relevance.
What's in the full article
Synack's full blog post covers the operational detail this post intentionally leaves for the source:
- Detailed build-vs-buy reasoning from the perspective of a security vendor operating AI pentesting at scale
- Discussion of token consumption, orchestration overhead and the economics of repeated agentic workflows
- Examples of how model deprecation and pricing shifts affect product maintenance over time
- The webinar reference point where Synack and Dow discussed production-grade AI pentesting decisions
👉 Read Synack's analysis of the build-vs-buy decision for AI pentesting →
AI pentesting at scale: can in-house teams sustain the economics?
Explore further
AI pentesting is becoming an identity governance problem as soon as the workflow depends on third-party models and service credentials. The article is about cost and reliability, but the control surface includes model APIs, orchestration identities and the human validation chain. That means teams need to think about who or what can invoke the testing workflow, how access is bounded, and how lifecycle changes are governed. In practice, AI security testing inherits the same access-control discipline as other sensitive automation.
A question worth separating out:
Q: When does AI pentesting become too costly to run in-house?
A: It becomes expensive when the team moves from occasional tests to repeated enterprise-scale usage. The cost drivers are not only model tokens and compute, but also orchestration, monitoring, human review and re-engineering when providers change the underlying model lifecycle.
👉 Read our full editorial: AI pentesting economics are the real build-vs-buy decision