Manual red teaming creates risk because it is point-in-time, labour intensive, and usually covers only a fraction of the environment. Threat actors, by contrast, search continuously and can pivot across exposed services, credentials, and human targets. When validation is infrequent, blind spots persist between exercises, and attackers need only one successful path to compromise.
Why Manual Red Teaming Leaves Gaps Between Exercises
Manual red teaming is valuable, but it is inherently episodic. A team may validate one set of assumptions, test a limited slice of the estate, and confirm a specific chain of compromise, yet the environment keeps changing after the assessment ends. New applications appear, cloud configurations drift, credentials rotate, third parties change integrations, and defenders patch one weakness while another remains open. The result is not that red teaming has no value, but that its findings decay quickly when they are not paired with continuous exposure management and detection improvement.
For that reason, the main issue is coverage and timing rather than competence. A manual exercise can show whether a pathway is feasible at a point in time, but it cannot continuously observe whether the same pathway, or a different one, becomes viable next week. That is why continuous validation matters, and why practitioners often compare red team findings with broader attack-surface management or adversary emulation evidence rather than treating a single engagement as proof of resilience. In practice, many security teams discover the longest-lived gaps only after an attacker has already used a quieter route than the one the red team explored.
When the exercise ends, the environment does not stop evolving, so the test result is always already aging.
How Continuous Exposure Differs From a One-Off Exercise
Manual red teaming answers a narrow but important question: can a skilled adversary achieve a defined objective under the current conditions? That makes it useful for validating assumptions, exercising defenders, and finding high-value weaknesses. It does not, however, behave like a persistent control. It is not designed to watch every asset, retest every change, or keep pace with the daily churn of identity, endpoint, and cloud exposure.
That difference matters because modern compromise paths are often multi-stage. Attackers may combine a public-facing weakness with credential abuse, weak segmentation, or phishing, and they can do this repeatedly until one path works. A manual team normally works within a limited scope and timeframe, so it may miss lower-visibility paths that only become attractive once the obvious route is fixed. This is also why external validation sources can be useful when they clarify the gap between a point-in-time assessment and adversary persistence, such as Anthropic’s report on an AI-orchestrated cyber espionage campaign, which illustrates how persistence and adaptation change the defender’s burden.
- A manual test is strongest when it validates a defined assumption, not when it is treated as continuous assurance.
- Coverage is usually bounded by budget, scope, access, and time, so the test may miss adjacent weaknesses that remain exploitable.
- Detection and response maturity often matter as much as initial exploitation, because failed containment can turn a partial foothold into an incident.
The model breaks down when organisations expect a one-time engagement to substitute for ongoing discovery, monitoring, and re-validation after material change.
Where the Assumptions Behind Manual Testing Break Down
Tighter hands-on testing often gives deeper insight into attack chains, but it also increases cost and reduces frequency, so organisations must balance depth against freshness. That tradeoff is real, and it is why there is no consensus that manual red teaming alone is the right answer for every environment. For high-churn estates, the most defensible approach is usually to use manual testing for depth and judgement, then pair it with more continuous forms of exposure tracking and control verification.
The edge cases are often where exposure lasts longest. Large hybrid environments change too quickly for periodic validation to remain complete. Third-party dependencies can create new paths that were absent during the last exercise. Identity abuse can also outlive technical fixes when standing access, stale privileges, or weak revocation processes remain in place after an engagement. Those are not failures of red teaming itself; they are failures of organisations to treat findings as part of an ongoing lifecycle.
Where teams most often go wrong is assuming that a successful hardening action closes the broader class of risk. A patched control may remove one exploit, but the same business process or trust boundary can still be reachable through another route. The safer interpretation is that red teaming proves exposure exists, while continuous validation proves whether the exposure has actually stayed closed.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
MITRE ATT&CK address the attack and risk surface, while CIS Controls v8 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.
| Framework | Control / Reference | Relevance |
|---|---|---|
| MITRE ATT&CK | T1589 — Gather Victim Identity Information | Manual red teams often emulate attacker reconnaissance and multi-stage compromise paths. |
| T1595 — Active Scanning | Point-in-time testing can miss continuously changing exposed services and attack surface. | |
| Recommendation — Map observed red-team paths to ATT&CK techniques and validate detections for each stage. Use active-scanning techniques to retest exposed assets between manual exercises. | ||
| CIS Controls v8 | CIS-01 — Inventory and Control of Enterprise Assets | Long-lived exposure often persists because asset scope changes faster than periodic testing. |
| CIS-07 — Continuous Vulnerability Management | Manual red teaming does not provide ongoing exposure verification after conditions change. | |
| Recommendation — Maintain current asset inventories so remediation can be revalidated against the full estate. Run continuous vulnerability management to detect reopened or newly introduced exposure. | ||
| NIST CSF 2.0 | DE.CM — Security Continuous Monitoring | The problem is stale assurance between assessments, which continuous monitoring addresses. |
| ID.AM — Asset Management | Periodic exercises miss weaknesses when asset scope and services change after the test. | |
| Recommendation — Implement continuous monitoring to shorten the time between exposure creation and detection. Keep asset scope current so red-team lessons apply to the actual production estate. | ||
Practitioner Guidance
What to prioritise: Treat red team findings as change triggers. The highest-value follow-up is to verify whether the same weakness still exists after remediation, whether adjacent systems share the same pattern, and whether detections now alert on the behaviour rather than only the exact technique.
What to verify: Confirm that scope drift is not hiding reintroduced exposure. If the environment changes faster than the next planned exercise, the security team should assume the last result is incomplete until the affected services, identities, and paths are rechecked.
Common mistake: Using red team success or failure as a binary measure of security posture. A failed engagement does not mean the environment is safe, and a successful engagement does not mean the specific path is the only concern.
Practitioner takeaway: Manual red teaming is best understood as a deep diagnostic, not a durable control; the organisations that reduce exposure fastest are the ones that convert every finding into continuous re-validation and detection improvement.
Related resources from NHI Mgmt Group
- Why do EDR alerts still leave organisations exposed if response is manual?
- Why do production scans and quarterly penetration tests leave organisations exposed for too long?
- Why does a once-a-year mobile penetration test leave organisations exposed for so long?
- When do short-lived access tokens still leave organisations exposed?