By NHI Mgmt Group Editorial TeamBased on ControlMonkey: “Affected by the AWS Outage? 5 Things to do Tomorrow for your Cloud Resilience” (October 20, 2025)

TL;DR: AWS’s recent outage triggered more than 6.5 million disruption reports worldwide and exposed a harder truth for cloud teams: disaster recovery fails when configuration, dependencies, and drift are not recoverable, according to ControlMonkey and CNN. Data backups alone do not restore operational identity, policy state, or infrastructure topology.


At a glance

What this is: This is an analysis of why cloud disaster recovery failed in the AWS outage: configuration, dependencies, and drift mattered as much as data availability.

Why it matters: It matters because IAM, PAM, and cloud security teams cannot treat recovery as a storage problem when operational identity, policy state, and infrastructure topology are part of the control plane.


Context

Cloud disaster recovery is the ability to restore running services, not just recover files. This article argues that the AWS outage exposed a familiar governance gap: teams often protect backups but not the configuration, dependencies, and region-specific assumptions that make workloads actually run.

For identity and infrastructure teams, the issue is broader than a storage restore. If policies, IaC state, service dependencies, and operational topology are not recoverable together, the environment may come back partially or with broken controls after a regional failure.


Key questions

Q: What breaks when cloud disaster recovery only restores data?

A: Recovery breaks when teams cannot reconstruct the configuration, permissions, and dependencies needed for workloads to run. A backup may restore files, but it does not automatically restore network paths, identity policy, or service topology. That leaves the environment partially restored and slower to recover.

Q: Why do configuration drift and manual changes increase cloud outage risk?

A: Because they create a gap between the environment you think you can restore and the environment that actually exists. When failover scripts, runbooks, or redeployments expect one configuration and production contains another, recovery becomes unpredictable and security controls can be missed or broken.

Q: How should security teams test whether cloud recovery actually works?

A: They should run full recovery exercises that rebuild the environment, not just restore data. The test should confirm IAM permissions, network paths, dependencies, and application behavior all work together under outage conditions. If a restored system cannot run as intended, the backup programme has only proven retention, not recovery.

Q: What should teams do immediately after finding gaps in infrastructure as code coverage?

A: Bring manual or hidden resources into source control, then reconcile the declared state with production before the next outage. Any component that only exists in a console or in tribal knowledge is a recovery liability because it cannot be reproduced consistently under pressure.


Technical breakdown

Why backups do not restore cloud operations

Traditional disaster recovery assumes that data loss is the main failure mode. In cloud environments, service availability also depends on infrastructure definitions, identity policies, network controls, regional placement, and service dependencies. If those elements were created manually, drifted from code, or never inventoried, restoring the workload can recreate the application without recreating the operating conditions it needs. That is why a backup can be intact while the service remains unusable. The recovery target is not only the data set but the full configuration state required for the workload to function.

Practical implication: Treat disaster recovery as a configuration recovery problem, not a storage-only exercise.

How configuration drift breaks failover and redeployment

Configuration drift occurs when production no longer matches the desired state defined in infrastructure as code or other source-of-truth controls. During an outage, that mismatch can break automated redeployments, produce inconsistent failover behaviour, or reintroduce security gaps that were not present in the intended design. The article’s point is practical: if manual console changes, hidden dependencies, or shadow resources exist, the recovery path becomes unpredictable. In cloud recovery, drift is not an administrative nuisance. It is a direct cause of failed restoration because the recovered environment no longer matches what the runbook expects.

Practical implication: Detect drift continuously and reconcile it before an outage turns it into a recovery failure.

Why dependency mapping is part of recovery design

Cloud recovery depends on understanding what a workload actually uses, not only what appears in the primary application stack. Regional placement, third-party services, IAM dependencies, and supporting tooling can all become single points of failure if they are not documented and recoverable. The article notes that many teams only learn their critical workloads are tied to a specific region when the region fails. That is a governance problem, not just an operations problem, because the recovery design did not model the true service graph. Without dependency mapping, failover plans miss the components that determine whether services can restart cleanly.

Practical implication: Map dependencies into the recovery plan so failover covers the full service graph, not only the application.


NHI Mgmt Group analysis

Configuration is now a recovery asset, not an implementation detail. This outage shows that cloud disaster recovery fails when teams treat infrastructure state as less important than data copies. In practice, the environment that comes back after an incident must include policy, network, identity, and dependency state, or recovery is only partial. For cloud governance, the recovery unit is the live configuration baseline, not the backup file.

Infrastructure as code coverage defines recoverability. Where teams still rely on console changes, ad hoc fixes, or untracked resources, disaster recovery becomes nondeterministic. The operational question is not whether the data can be restored, but whether the environment can be reproduced with the same permissions, dependencies, and routing assumptions. That is why IaC coverage is a resilience control, not only a delivery practice.

Drift is the hidden failure mode behind many cloud outages. The outage illustrates a named concept worth tracking: configuration recovery gap. When desired state and actual state diverge, the recovery script may succeed technically while the service still fails operationally. Practitioners should treat drift as a first-class recovery risk because it undermines both repeatability and assurance.

Cloud resilience has to extend across provider and third-party boundaries. The article correctly frames the issue beyond AWS itself. Modern workloads depend on cloud providers, observability platforms, and regional services that can all affect restoration outcomes. Teams that scope recovery only to the primary cloud account are underestimating the blast radius of an outage.

Operational identity is part of disaster recovery scope. When restoration depends on policies, permissions, service accounts, and environment-specific dependencies, IAM governance becomes part of the recovery design. That means resilience reviews should cover identity state alongside compute and storage, because a restored workload that cannot authenticate or authorize is not recovered.

What this signals

Configuration recovery gap: Many cloud programmes still assume that a good backup means a recoverable service. This outage shows the real boundary is broader: policy state, dependencies, and infrastructure topology all have to be restorable together.

Teams should treat drift detection and IaC coverage as resilience controls rather than delivery hygiene. If the environment cannot be reproduced deterministically, the recovery process will inherit the same uncertainty that caused the outage to hurt in the first place.


For practitioners

  • Audit the full recovery surface Inventory services, regions, dependencies, and shadow resources so the recovery plan reflects what workloads actually require to run.
  • Close infrastructure as code gaps Move legacy stacks, ClickOps-created resources, and manual configuration into code so recovery can be reproduced deterministically.
  • Test regional failover with mini drills Simulate a single-region outage for one critical service and measure whether runbooks, automation, and dependencies restore the service end to end.
  • Detect and remediate configuration drift Use automated drift detection to keep production aligned with declared state and to prevent recovery mismatches during restoration.
  • Snapshot infrastructure and policy state daily Capture configuration, dependencies, and policies together so recovery can restore the environment, not just the data.

Key takeaways

  • Cloud disaster recovery fails when teams restore data without restoring the configuration that makes applications operable.
  • The outage exposed a wider resilience problem because dependencies, policies, and regional assumptions were part of the failure surface.
  • Practitioners need continuous drift detection, stronger IaC coverage, and recovery testing that proves the full environment can come back.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0PR.IR-01 — Network ResilienceThe article centres on recovery of cloud services after regional failure.
PR.DS-11 — Data BackupBackups are discussed as necessary but insufficient for restoring operations.
Recommendation — Align recovery design to PR.IR-01 so services can be restored after a regional outage. Use PR.DS-11 to ensure backup strategy supports recovery of both data and configuration.
NIST SP 800-53 Rev 5CP-9 — System BackupThe article contrasts data backup with full cloud recovery.
CP-10 — System Recovery and ReconstitutionRestoration of the operating environment is the core subject here.
Recommendation — Apply CP-9 alongside configuration recovery so backups do not stop at files alone. Use CP-10 to verify reconstitution of the full cloud environment, not just data.
CIS Controls v8CIS-4 — Secure Configuration of Enterprise Assets and SoftwareConfiguration drift and untracked resources drive the recovery failure described.
Recommendation — Use CIS-4 to keep production configuration aligned with the recoverable baseline.

Key terms

  • Configuration Recovery: Configuration recovery is the ability to restore the cloud settings, access controls, and service dependencies needed to make applications run again. It goes beyond data backup by preserving the operating state of infrastructure, identity, network, and security components so teams can return to a known-good environment after an outage or change.
  • Configuration Drift: Configuration drift is the gradual divergence between a system's intended secure state and the settings it actually runs with over time. In SaaS, drift often appears when admins change sharing, logging, or access controls under pressure and never return to validate the result.
  • Infrastructure as Code coverage: The share of infrastructure that is created, changed, and governed through code rather than manual console actions. In practice, it measures how much of the environment can be reviewed, reproduced, and remediated through a controlled delivery path instead of ad hoc operator behaviour.
  • Recovery Dependency Mapping: A structured view of which systems, identities, credentials, and data flows must return before a service can be considered functional. In practice, it helps teams sequence restoration so access control and business operations come back together rather than separately.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 11, 2026.
Updated on October 10, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org