Join our Newsletter — 33% off our NHI Course

Who is accountable for testing recovery plans before a ransomware event exposes gaps in resilience?

Security, identity, and infrastructure owners all share accountability, but governance should assign clear ownership for recovery testing, evidence capture, and remediation follow-up. Recovery plans that are only documented, not exercised, fail under pressure. Organisations should treat restoration testing as an operational control, with named owners, recurring validation, and board-visible reporting on readiness.

Why This Matters for Security Teams

Recovery testing is not a paperwork exercise. When ransomware hits, the failure is rarely the existence of a plan and more often the absence of proof that the plan actually works under pressure. Security teams, identity teams, and infrastructure owners all touch restoration, but accountability must be explicit because gaps in backup integrity, credential recovery, and rebuild sequencing often sit between functions. NIST’s Cybersecurity Framework 2.0 treats resilience as an operational outcome, not a document shelf item.

NHIMG research shows why this matters: in the Ultimate Guide to NHIs, 91.6% of secrets remain valid five days after notification, which means recovery and revocation often lag long after an incident is known. That same delay can undermine restoration paths if service accounts, API keys, or automation tokens are not tested as part of the recovery workflow. In practice, many security teams discover broken restore steps only after attackers have already exposed the weakest dependency chain, rather than through intentional resilience testing.

How It Works in Practice

Accountability should be assigned to the control owner who can prove end-to-end recovery, not just the team that stores the backup. In most organisations, that means security owns the testing standard, infrastructure owns platform restoration, and identity owners verify privileged access, secrets, and recovery accounts are usable during rebuild. Current guidance suggests treating recovery exercises like any other control: define scope, assign an owner, require evidence, and track remediation to closure.

A practical recovery test should include:

  • Restoring data from immutable or offline backups and validating integrity before production use.
  • Rebuilding identity dependencies such as directory services, service accounts, and break-glass access.
  • Testing revocation and re-issuance of secrets so compromised credentials do not survive the reset.
  • Measuring recovery time objectives against real steps, not aspirational estimates.
  • Capturing evidence for audit, risk, and board reporting, including failures and follow-up actions.

This is especially important for environments where automation depends on non-human identities. NHIMG’s 52 NHI Breaches Analysis and the MGM Resorts Breach 2023 both show how identity compromise can turn operational recovery into a second incident if restoration is not coordinated with access control. These controls tend to break down when rebuilds are staged across multiple cloud accounts and legacy directories because ownership boundaries make it unclear who can prove the system is actually recoverable.

Common Variations and Edge Cases

Tighter recovery testing often increases operational overhead, requiring organisations to balance assurance against maintenance windows, change freezes, and system complexity. That tradeoff is real, but it does not remove the need for testing. Best practice is evolving toward risk-based cadence, where the most critical business services, privileged identities, and externally exposed systems are exercised more frequently than low-impact workloads.

There is no universal standard for test frequency, but current guidance suggests a mixed model: tabletop validation for coordination, technical restore tests for data and infrastructure, and identity recovery tests for accounts, secrets, and access pathways. The NIST SP 800-53 Rev. 5 Security and Privacy Controls reinforces that contingency and recovery capabilities should be testable, while the ENISA Threat Landscape continues to highlight ransomware as an operational resilience issue, not only a malware issue. The practical edge case is third-party managed recovery: if a vendor controls backups, identity federation, or failover tooling, the organisation still needs named internal accountability for evidence, escalation, and remediation. Without that, recovery becomes a shared assumption instead of a tested capability.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 and CSA MAESTRO address the attack and risk surface, while NIST CSF 2.0, NIST AI RMF and NIST SP 800-63 set the governance and control requirements practitioners need to meet.

Framework Control / Reference Relevance
NIST CSF 2.0 RC.RP Recovery planning and testing are core resilience functions.
OWASP Non-Human Identity Top 10 NHI-07 Covers recovery of non-human identities and secrets after compromise.
CSA MAESTRO GOV-03 Governance must define responsibility for agent and workload recovery assurance.
NIST AI RMF GOV-4 Governance requires accountability for reliable operation and recovery readiness.
NIST SP 800-63 Recovery processes must preserve identity assurance during account restoration.

Test secret rotation, revocation, and service account recovery as part of every restoration drill.