Join our Newsletter — 33% off our NHI Course

How should security teams design for operational continuity during an attack?

Design around the systems that must stay live, then segment them into smaller trust zones with explicit isolation paths. Continuous containment matters more than perfect prevention, because the control objective is to keep critical operations running while pressure is applied.

Designing continuity around the most critical live systems

Operational continuity starts by identifying which services must keep working under hostile conditions, then designing the environment so those services have smaller blast radiuses, clear dependencies, and explicit isolation paths. The point is not to preserve every system equally. It is to keep essential operations available while the rest of the environment absorbs, contains, and recovers from pressure.

This is why continuity planning cannot be only a recovery exercise. If the “must stay live” set is not defined in advance, teams tend to overprotect low-value assets and underprotect the paths that actually keep the business running. For continuity during an attack, architecture becomes a resilience control, not just a security design choice.

NIST Cybersecurity Framework 2.0 is useful here because it ties continuity to the full cycle of identifying essential services, protecting them, detecting disruption, responding to it, and recovering without losing control of the environment.

Why segmentation must support containment, not just separation

Segmentation only helps if it is designed for partial failure. Smaller trust zones, constrained east-west paths, and explicit allowlists reduce the chance that an intrusion in one area becomes a company-wide outage. In practice, that means the network, workload, and access boundaries around critical functions should be tighter than the boundaries around ordinary business systems.

Continuity-focused segmentation also has to assume that some controls will fail. If teams rely on a single policy layer, a single management plane, or a single identity path to preserve isolation, they create a new point of failure. Good design gives responders room to contain the attack without breaking the operational path that supports the core service.

Protect and respond functions in NIST CSF 2.0 help frame this trade-off: the design goal is bounded degradation, not perfect prevention.

How to keep control when pressure is already underway

When an attack is active, the continuity question shifts from “how do we stop everything?” to “what can we safely disconnect, reduce, or reroute without collapsing operations?” That requires preplanned failover, manual fallback paths for critical workflows, and clear ownership for emergency decisions. It also requires evidence of which dependencies are truly essential, because dependency maps that look complete in calm conditions often miss the hidden services that become critical during incident response.

Teams should be especially careful with privileged administrative paths, management interfaces, and shared services. Those are often the routes attackers target to turn an isolated compromise into operational disruption. Keeping continuity means protecting the control plane as carefully as the production workload, and knowing when to degrade functionality deliberately rather than let the attacker decide the failure mode.

NIST SP 800-207 Zero Trust Architecture reinforces the design logic behind explicit trust zones, while MITRE ATT&CK Enterprise is useful for anticipating the attack paths that most often threaten continuity, such as credential access, lateral movement, and privilege escalation.

Risk and Threat Considerations

The main continuity risk is that organisations discover too late that their “resilient” design still depends on a few shared services, shared credentials, or shared administrative paths. Once those dependencies are attacked, a local incident can become a wide outage, or the response itself can take down the live path that the business needs most.

Failure mechanism: Attackers often exploit excessive trust between zones, shared management channels, or overbroad access to move from an initial foothold into the systems that carry core operations. If containment is weak, defenders may be forced to choose between shutting down the environment and leaving the attacker in place.

Impact: The business loses the ability to control degradation. Instead of an intentional, bounded interruption, the organisation gets cascading failure, longer recovery, and potentially a complete loss of confidence in the live service path.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0 provides the primary governance reference for this topic.

Framework Control / Reference Relevance
NIST CSF 2.0 PR.AA-05 — Access Permissions and Authorizations Are Managed Continuity depends on limiting who can reach critical live systems during attack.
RC.RP-01 — Recovery Plan is Executed During or After an Incident The question is about keeping essential operations running under attack and then restoring control.
DE.CM-01 — Networks and Network Services Are Monitored to Find Potential Cybersecurity Events Continuity requires visibility into attack pressure on the paths that keep live services running.
Recommendation — Restrict administrative and service access paths to preserve containment and operational control. Practice recovery playbooks that preserve essential service availability under hostile conditions. Monitor critical trust zones and service paths for signs of containment failure.

Practitioner Guidance

What to prioritise: Define the minimum viable set of services that must remain available, then design their dependencies, communications, and administrative paths as if they will be actively contested.

What to verify: Test whether a defender can isolate a compromised zone without taking down the critical service itself. If the answer is no, the continuity plan is still too coupled.

Decision rule: If a control improves prevention but makes the live service easier to disrupt during response, prefer the control design that preserves containment and operational control first.

Practitioner takeaway: Continuity during attack is won by designing for controlled degradation, not by assuming prevention will hold long enough for recovery to be easy.