They fail when organisations mistake data availability for service recoverability. In hybrid and multi-cloud estates, the harder problem is restoring the complete set of dependencies, identities and trust relationships that make a service usable. Tools can store copies and document plans, but they cannot guarantee that the restored environment will function in a real incident.
Why backup jobs can succeed while recovery still fails
Backup and disaster recovery tools often optimise for copy creation, retention, and restore mechanics, but complex environments fail at a different layer. The real test is whether the restored workload can authenticate, resolve dependencies, reach adjacent services, and satisfy policy requirements after failover. When those conditions are broken, a successful restore becomes only a partial recovery.
What complexity changes in hybrid and multi-cloud recovery
In simpler estates, a backup set may map cleanly back to one application, one platform, and one control plane. In hybrid and multi-cloud environments, the service is usually spread across identity providers, DNS, certificates, network controls, managed services, queues, caches, and external integrations. Recovery must therefore rebuild the service context, not just the files or virtual machines.
That is why service dependencies matter as much as data copies. A database snapshot is useless if the application cannot rebind to it, a cluster cannot rejoin the control plane, or the restored environment still points to stale endpoints and expired secrets. The more distributed the architecture, the more recovery becomes a systems-integration problem rather than a storage problem.
Why “restore” is not the same as “operationally usable”
Tools commonly assume that the restore target resembles the source environment at the time of backup. In practice, dependencies drift, permissions change, certificates expire, and cloud services evolve. That gap means the restore can complete while the application remains broken, degraded, or non-compliant.
Recoverability also depends on trust. If identities, access paths, and service-to-service relationships are not restored in the right order, the rebuilt system may fail closed or, worse, come back with excessive access just to make the application run. Both outcomes are a recovery failure, because one blocks service and the other creates exposure.
Risk and Threat Considerations
The main risk is false confidence. Organisations can believe they have a working recovery capability because backups exist and restore tests pass in narrow conditions, while the real incident path fails on dependency order, access, or trust reconstruction. In an outage or cyber event, that gap extends downtime and can force unsafe workarounds.
Failure mechanism: Recovery breaks when the tool restores data or infrastructure without re-establishing the surrounding service dependencies, identity bindings, and trust relationships in a valid sequence. The environment may be intact on paper but unusable in practice.
Impact: Recovery time increases, service restoration becomes manual and error-prone, and teams may reintroduce access, network, or configuration shortcuts that enlarge blast radius during the incident.
Standards & Framework Alignment
This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.
NIST CSF 2.0, NIST SP 800-53 Rev 5 and CIS Controls v8 set the technical controls, while ISO/IEC 27001:2022 defines the regulatory obligations.
| Framework | Control / Reference | Relevance |
|---|---|---|
| NIST CSF 2.0 | RC.RP-01 — Recovery Plan Execution | Recovery must restore the full service path, not just data copies. |
| ID.AM-03 — Asset Management: Organizational communication and data flows are identified and managed | Complex recovery depends on knowing service dependencies and data flows. | |
| PR.AA-05 — Identity Management, Authentication and Access Control | Recovered services must re-establish authentication and access relationships. | |
| Recommendation — Validate that recovery plans restore the service and its dependencies, not only the backup artifact. Map application dependencies and data flows so recovery can be sequenced correctly. Verify that recovered systems can authenticate and authorize users and services before declaring recovery complete. | ||
| NIST SP 800-53 Rev 5 | CP-10 — System Recovery and Reconstitution | This control directly covers restoring systems to an operational state after disruption. |
| CP-2 — Contingency Plan | Contingency planning must account for complex restoration dependencies and conditions. | |
| IA-5 — Authenticator Management | Recovery often fails when secrets, keys, and authenticators are expired or unusable. | |
| Recommendation — Exercise full reconstitution steps, including dependencies and access prerequisites, during recovery testing. Document contingency procedures that include sequencing, prerequisites, and service validation. Track and validate authenticators, keys, and secrets as part of recovery readiness. | ||
| ISO/IEC 27001:2022 | A.5.30 — ICT readiness for business continuity | Business continuity readiness includes the ability to restore ICT services for use. |
| A.8.13 — Information backup | Backups are necessary but insufficient without recovery validation. | |
| Recommendation — Confirm that continuity plans restore ICT services to an operationally usable state. Test that backups can be restored into a working service, not just retained securely. | ||
| CIS Controls v8 | CIS-11 — Data Recovery | Recovery controls must prove recoverability, not only data preservation. |
| Recommendation — Regularly test restores against real application and dependency requirements. | ||
Practitioner Guidance
What to verify: Test recovery against the whole service path, not just the backup artifact. A valid test should prove that the application starts, authenticates, reaches its dependencies, and supports the intended business transaction after restoration.
What good looks like: Recovery plans are dependency-aware, version-aware, and environment-aware. Teams can show that the same runbook works when secrets, certificates, network policies, and cloud service bindings are rotated or rebuilt.
Decision rule: If a restore test only confirms that data returned, treat it as a backup test, not a recovery test. If the service cannot be exercised end-to-end, the control is not yet proving recoverability.
Practitioner takeaway: The standard for disaster recovery is not whether copies exist, but whether the restored service can operate safely and predictably in the real dependency graph it actually depends on.
Related resources from NHI Mgmt Group
- Why do legacy IGA tools fail to control over-entitlement and access drift in complex environments?
- Why do complex passwords still fail in real environments?
- Why do native self-service reset tools fail more often in hybrid environments?
- Why do access certification programmes fail in complex environments?
Deepen Your Knowledge
Free weekly newsletter
Subscribe to the NHI & AI Identity Journal
The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.
Bonus 33% off our NHI Course when you subscribe.
Reviewed and updated by the NHIMG editorial team on October 11, 2026.
NHI Mgmt Group — the #1 independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org