By NHI Mgmt Group Editorial TeamBased on Oasis Security: “Don’t Look Back In Anger: How Cloudflare’s Outage Highlights the Need for Safer Rotations” (May 1, 2026)

TL;DR: Cloudflare’s March 21, 2025 outage lasted 1 hour and 7 minutes, caused total write failures and partial read failures in R2, and stemmed from a credential rotation error that updated the wrong environment before old keys were deleted, according to Oasis Security. The lesson is structural: rotation without verification turns NHI lifecycle control into an outage trigger, not a safeguard.


At a glance

What this is: This is an analysis of Cloudflare’s March 21, 2025 R2 outage, where a credential rotation error triggered global write failures and degraded reads after the wrong environment was updated and old keys were removed.

Why it matters: It matters because NHI rotation is not just a security task but an availability control, and IAM teams need verification, inventory, and staged revocation to avoid turning lifecycle work into production downtime.

By the numbers:

  • Cloudflare’s March 21, 2025 outage lasted 1 hour and 7 minutes.

Context

Cloudflare’s outage shows that NHI rotation is not only about shortening secret lifetime. It is also about proving which environment received the new credential, confirming which identity is still active, and preventing revocation from outrunning validation.

For IAM, PAM, and NHI programmes, the failure mode is familiar: lifecycle controls that look safe on paper can become outage generators when inventory, ownership, and runtime verification are incomplete. The operational question is whether rotation is governed as a controlled change or treated as a simple replacement task.


Key questions

Q: What fails when NHI rotation is done without verification?

A: The cutover can succeed in one environment while production still depends on the old credential, so revocation turns into an outage. Safe rotation requires proof of active use in the target environment before deletion of the prior secret.

Q: When does credential rotation create more risk than it reduces?

A: Rotation becomes risky when teams do not understand which applications depend on a credential or how widely it is used. In those cases, blind rotation can disrupt production and make security teams hesitate to act again. The answer is not to avoid rotation, but to map dependencies first and automate change control.

Q: What are the signs that NHI rotation controls are too weak?

A: Repeated environment mismatches, manual cutovers, missing ownership metadata, and revocations that cannot be reversed are strong indicators. Those signals show the programme is assuming success instead of proving it before decommissioning the old credential.

Q: What should teams do when a credential rotation is about to retire an active secret?

A: Keep the old credential in place until the new one is verified in production, then retire the prior secret under change control. If the system cannot prove active use and ownership, the rotation should not proceed to deletion.


Technical breakdown

Why credential rotation fails without environment verification

Credential rotation is meant to replace one authenticator with another without breaking the identity relationship in production. In this case, the new key pair was generated correctly, but the deployment landed in the default development environment because the production flag was omitted. That created a split-brain state: one environment believed rotation had succeeded, while production still depended on the old secret. The underlying issue is not rotation itself, but the lack of a verified handoff between issuance, deployment, and cutover.

Practical implication: require environment-specific verification before any old credential is revoked.

Why NHI lifecycle management needs live usage context

A secret is safe to retire only when the system can prove the new credential is actually in use and the old one is idle. Cloudflare’s outage shows what happens when teams rely on process assumption instead of runtime evidence. Without a clear map of where each secret is consumed, operators cannot distinguish a successful cutover from a partial one. That makes deletion an availability risk, not just a security action. This is a lifecycle governance problem as much as a technical one.

Practical implication: tie rotation approval to consumption evidence, ownership metadata, and dependency mapping.

How manual rotation turns into an outage path

Manual credential handling increases the chance that one missed parameter or one false assumption will cascade into service failure. Here, the wrong environment was updated, old keys were deleted, and production continued using invalid credentials. The result was total write failure and partial read degradation in R2. This is the classic NHI failure pattern where governance is present as a policy, but not as an enforced control path. The architecture lacked guardrails strong enough to stop a premature revocation.

Practical implication: move from manual rotation steps to staged, policy-guarded rotation with rollback protection.


Threat narrative

Attacker objective: There was no external attacker objective in the article’s primary incident; the objective was intended credential replacement, but the control failure produced service outage instead.

  1. Entry occurred through a rotation change that generated a new credential pair for the R2 Gateway.
  2. Credential abuse followed when the new credential was deployed to the wrong environment and production remained bound to the old secret.
  3. Escalation took the form of operational failure as the backend revoked the still-active credential while production traffic continued to use it.
  4. Impact was global service disruption, with total write failures and partial read failures in Cloudflare R2 for 1 hour and 7 minutes.
  • Okta support system breach 2023: A support service account credential saved in a personal Google profile let attackers take HAR files and hijack five Okta customers' sessions.
  • Dropbox Sign breach 2024: A compromised back-end service account gave attackers Dropbox Sign customer data, including API keys, OAuth tokens and MFA information.

Read our 52 NHI Breaches Analysis report for a comprehensive view of breaches impacting Non-Human Identities including AI Agents.


NHI Mgmt Group analysis

Rotation without verification is not a security control, it is an outage trigger. Cloudflare’s incident shows that replacing a secret is only safe when the organisation can prove the new credential reached the intended environment and the old one is no longer active. The control failed because revocation was treated as a routine follow-on, not a gated decision. For NHI governance, the implication is that cutover evidence must be part of the control, not an optional audit trail.

Identity lifecycle for NHIs breaks when ownership and runtime context are missing. A secret can only be retired safely if teams know what owns it, where it is consumed, and what dependency graph depends on it. Cloudflare’s outage reflects a familiar NHI governance gap: the identity was managed as a credential object, not as a live production dependency. The practitioner conclusion is that lifecycle policy without consumption context is incomplete governance.

Safe rotation depends on staged change, not instantaneous replacement. The incident reinforces that NHI lifecycle management should be treated like any other high-risk production change, with validation before decommissioning and rollback available until the new path is proven. That does not make rotation slower for its own sake; it makes the control usable in real environments. Practitioners should judge rotation maturity by whether revocation is reversible until verification is complete.

Cloudflare’s outage validates a broader named concept: rotation blast radius. The blast radius is the operational damage created when secret replacement touches active dependencies before the new state is confirmed. In this case, the blast radius was write failure, read degradation, and global service impact. The practitioner conclusion is that rotation policy must be designed around containment of failure, not just reduction of secret age.

This incident fits OWASP-NHI and NIST lifecycle thinking more than a narrow key-management view. The problem was not simply that a key existed, but that the lifecycle process could not distinguish safe retirement from unsafe revocation. That is why inventory, offboarding, and verification belong in the same governance conversation. The practitioner takeaway is to treat NHI rotation as governed change management, not a standalone cryptographic task.

From our research library:

What this signals

Rotation blast radius: The Cloudflare incident shows that the real risk is not secret age alone, but the operational blast radius created when revocation outruns verification. NHI programmes should assume that every rotation can fail unless the new credential is proven active in the live environment first.

Lifecycle governance has to include runtime evidence. Ownership metadata, dependency mapping, and consumption context are the controls that prevent a credential change from becoming an availability event. Without them, teams are managing objects, not identities.

The broader signal for practitioners is that secrets management and change management are now the same operational problem in production environments. A rotation process that cannot tell you where a credential is live is not mature enough for critical services.


For practitioners

  • Audit production cutover checks Require a positive verification step that confirms the new secret is active in production before any old key is deleted.
  • Map credential ownership and runtime dependencies Tag every service account, token, and backend secret with owner, environment, and consuming service so revocation decisions are not guesswork.
  • Stage rotations with rollback protection Use phased rotation so the old credential remains available until the new path is validated and the service dependency map is confirmed.
  • Separate production and non-production controls Prevent default-environment drift by making production deployment explicit and blocking deletion if the credential has not been observed in use.

Key takeaways

  • Cloudflare’s outage was caused by a credential rotation mistake, not by a shortage of security policy, which is why verification matters as much as rotation itself.
  • The incident produced total write failures and partial read failures in R2 for 1 hour and 7 minutes, showing that a small identity change can become a service-wide event.
  • The control gap was premature revocation without proof of production cutover, so teams need staged rotation, dependency mapping, and live-use verification before deleting old secrets.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

OWASP Non-Human Identity Top 10 addresses the attack and risk surface, while NIST SP 800-53 Rev 5, CIS Controls v8 and NIST CSF 2.0 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
OWASP Non-Human Identity Top 10NHI-01 — Improper OffboardingOld keys were deleted before production had fully cut over, creating an unsafe offboarding event.
NHI-07 — Long-Lived SecretsThe article centers on rotating secrets safely and avoiding stale credential persistence.
NHI-05 — Overprivileged NHIThe incident highlights the risk created when production access persists beyond safe control boundaries.
Recommendation — Treat credential retirement as offboarding and block deletion until production cutover is verified. Reduce standing secret lifetime and require verified replacement before retiring the old credential. Review NHI access scope so revocation cannot disrupt services that still depend on the old identity.
NIST SP 800-53 Rev 5IA-5 — Authenticator ManagementCredential rotation and revocation are directly governed by authenticator lifecycle management.
Recommendation — Apply IA-5 to enforce verified rotation, replacement, and revocation of production authenticators.
CIS Controls v8CIS-5 — Account ManagementThe outage stems from unsafe account and credential lifecycle handling across environments.
Recommendation — Use account management controls to inventory credentials and prevent premature deactivation.
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsThe failure occurred when authorization state changed without confirming active entitlement use.
Recommendation — Validate entitlements before revoking access and tie authorization changes to observed production use.

Key terms

  • Rotation blast radius: The operational impact created when a credential rotation affects active production dependencies before the replacement is confirmed. In NHI governance, blast radius measures how far a failed cutover can spread across services, environments, and teams before the error is contained.
  • Certificate Cutover: Certificate cutover is the point where traffic moves from one trust configuration to another, usually during a migration or infrastructure change. If the new certificate is not ready before DNS or routing changes take effect, clients may treat the service as insecure or untrusted.
  • Consumption context: The runtime information showing where a secret, token, or certificate is actually used. For NHIs, consumption context is what lets teams distinguish a live credential from an orphaned one, and it is essential for safe offboarding and revocation decisions.
  • Identity Dependency Mapping: Identity dependency mapping is the process of tracing which accounts, groups, sync flows, and applications rely on one another to function. It is essential in hybrid estates because a seemingly inactive identity may still support production access. Without it, lifecycle actions can break services or leave hidden privilege in place.

Deepen your knowledge

NHI governance, identity lifecycle management, and secrets management are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are building or maturing an IAM or NHI governance programme, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 6, 2026.
Updated on October 6, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org