TL;DR: Cloudflare’s March 21, 2025 outage lasted 1 hour and 7 minutes, caused total write failures and partial read failures in R2, and stemmed from a credential rotation error that updated the wrong environment before old keys were deleted, according to Oasis Security. The lesson is structural: rotation without verification turns NHI lifecycle control into an outage trigger, not a safeguard.
Editorial analysis by NHI Mgmt Group, based on content published by Oasis Security: “Don’t Look Back In Anger: How Cloudflare’s Outage Highlights the Need for Safer Rotations”.
By the numbers:
- Cloudflare’s March 21, 2025 outage lasted 1 hour and 7 minutes.
Key questions
Q: What fails when NHI rotation is done without verification?
A: The cutover can succeed in one environment while production still depends on the old credential, so revocation turns into an outage.
Q: When does credential rotation create more risk than it reduces?
A: Rotation becomes risky when teams do not understand which applications depend on a credential or how widely it is used.
Q: What are the signs that NHI rotation controls are too weak?
A: Repeated environment mismatches, manual cutovers, missing ownership metadata, and revocations that cannot be reversed are strong indicators.
Practitioner guidance
- Audit production cutover checks Require a positive verification step that confirms the new secret is active in production before any old key is deleted.
- Map credential ownership and runtime dependencies Tag every service account, token, and backend secret with owner, environment, and consuming service so revocation decisions are not guesswork.
- Stage rotations with rollback protection Use phased rotation so the old credential remains available until the new path is validated and the service dependency map is confirmed.
Bottom line: Cloudflare’s outage was caused by a credential rotation mistake, not by a shortage of security policy, which is why verification matters as much as rotation itself.
Explore further
View Full Forum → | NHI Foundation Course → | Our Services → | Read the full analysis →
Rotation without verification is not a security control, it is an outage trigger. Cloudflare’s incident shows that replacing a secret is only safe when the organisation can prove the new credential reached the intended environment and the old one is no longer active. The control failed because revocation was treated as a routine follow-on, not a gated decision. For NHI governance, the implication is that cutover evidence must be part of the control, not an optional audit trail.
A few things that frame the scale:
- 71% of NHIs are not rotated within recommended time frames, increasing the risk of compromise over time, according to the Ultimate Guide to NHIs.
A question worth separating out:
Q: What should teams do when a credential rotation is about to retire an active secret?
A: Keep the old credential in place until the new one is verified in production, then retire the prior secret under change control. If the system cannot prove active use and ownership, the rotation should not proceed to deletion.
👉 Read our full editorial: Cloudflare’s rotation outage shows why safer NHI rotations matter