Join our Newsletter — 33% off our NHI Course

Identity service outages: what IAM teams should do differently

 

(@nhi-mgmt-group)
Member Moderator
Joined: 1 year ago
Posts: 20739
Topic starter  

TL;DR: Regional AWS failure and cascading third-party outages disrupted Sign-On, AuthKit, and related services, with some request failure rates reaching 100% and roughly 70% of requests failing at peak during the second incident window, according to WorkOS. The lesson is that identity availability now depends on multi-layer resiliency, not just authentication correctness.

Editorial analysis by NHI Mgmt Group, based on content published by WorkOS: “Service disruption on October 20, 2025”.

Key questions

Q: What breaks when identity services depend on a single cloud region?

A: When identity services depend on a single cloud region, the login path can fail even if the rest of the application is intact.

Q: Why do identity outages matter even when existing sessions still work?

A: Because session continuity can mask the fact that the front door is closed.

Q: How should teams test resilience for hosted sign-in flows?

A: Teams should test the full sign-in path under dependency loss, not just the authentication backend in isolation.

Practitioner guidance

  • Map identity dependency chains Inventory every upstream service that a sign-in, session refresh, or admin access flow depends on, including hosting, flags, database connectivity, and edge delivery.
  • Test graceful degradation paths Simulate unavailability of feature flags, database lookup, and hosted page rendering to confirm that the identity experience times out cleanly rather than hanging.
  • Separate session continuity from sign-in availability Track whether existing sessions continue, whether new authentication works, and whether management pages render normally during dependency outages.

Bottom line: Identity outages can block access even when authentication rules and sessions are functioning as designed.

Explore further

View Full Forum →  |  NHI Foundation Course →  |  Our Services →  |  Read the full analysis →


This topic was modified 4 days ago by NHI Mgmt Group

   
Quote
(@mr-nhi)
Member Moderator
Joined: 5 months ago
Posts: 21545
 

Identity availability is now a control-plane governance problem, not just an infrastructure problem. This incident shows that authentication correctness does not preserve service continuity when credential lookup, hosting, and feature delivery all sit in the same failure domain. IAM teams should treat identity service resilience as part of the access architecture, because the user experience collapses before policy enforcement ever gets a chance to operate.

A question worth separating out:

Q: What is the difference between authentication correctness and identity service resilience?

A: Authentication correctness means the access decision is right when the service is available. Identity service resilience means the whole experience still functions when upstream components fail or degrade. In practice, a system can authenticate correctly in theory and still become unusable because page loads, credential retrieval, or hosting dependencies collapse.

👉 Read our full editorial: WorkOS outage shows why identity services need resilience beyond IAM


This post was modified 4 days ago by NHI Mgmt Group

   
ReplyQuote
Share:

Free weekly newsletter

Subscribe to the NHI & AI Identity Journal

The latest on NHI and Agentic AI security – articles, research, breaches, news and events every week.

Bonus 33% off our NHI Course when you subscribe.