By NHI Mgmt Group Editorial TeamBased on WorkOS: “Service disruption on October 20, 2025” (October 23, 2025)

TL;DR: Regional AWS failure and cascading third-party outages disrupted Sign-On, AuthKit, and related services, with some request failure rates reaching 100% and roughly 70% of requests failing at peak during the second incident window, according to WorkOS. The lesson is that identity availability now depends on multi-layer resiliency, not just authentication correctness.


At a glance

What this is: WorkOS describes how two outage windows across AWS and third-party dependencies disrupted identity services, especially AuthKit and Single Sign-On, despite some sessions remaining intact.

Why it matters: IAM teams need to treat identity availability as a resilience problem, because authentication correctness does not protect users if the surrounding service chain cannot fail over cleanly.


Context

Identity service resilience means the ability to keep sign-in, session handling, and related control-plane functions available when one dependency fails. This article shows why IAM programmes cannot stop at authentication correctness or access policy design if the service itself can no longer render pages or issue responses.

WorkOS experienced two outage periods on October 20 and 21, first from an AWS us-east-1 disruption and then from cascading failures at providers used for hosting and feature flags. The pattern is familiar in modern identity stacks: the identity plane is only as available as its weakest upstream service, and that makes operational resilience part of IAM governance.


Key questions

Q: What breaks when identity services depend on a single cloud region?

A: When identity services depend on a single cloud region, the login path can fail even if the rest of the application is intact. New connections, credential retrieval, and backend health checks may collapse together, which stops fresh authentication while leaving existing sessions partially unaffected. The result is a brittle access layer that cannot absorb regional outage without service interruption.

Q: Why do identity outages matter even when existing sessions still work?

A: Because session continuity can mask the fact that the front door is closed. Users may stay signed in while new logins, refreshes, or administrative actions fail. That split state creates a false sense of resilience and can hide a serious operational problem until business-critical access is already disrupted.

Q: How should teams test resilience for hosted sign-in flows?

A: Teams should test the full sign-in path under dependency loss, not just the authentication backend in isolation. That means simulating timeouts in database access, hosted page delivery, feature flags, and fallback logic. A resilient design keeps core authentication usable even when one upstream service becomes slow or unavailable.

Q: What is the difference between authentication correctness and identity service resilience?

A: Authentication correctness means the access decision is right when the service is available. Identity service resilience means the whole experience still functions when upstream components fail or degrade. In practice, a system can authenticate correctly in theory and still become unusable because page loads, credential retrieval, or hosting dependencies collapse.


Technical breakdown

Why identity services fail even when authentication is correct

Identity services often span several dependencies: application code, database connectivity, hosting, feature flags, and edge delivery. If any one of those layers is unavailable, authentication can be technically correct while the user-facing sign-in flow still fails. In this incident, a database proxy could not retrieve credentials because IAM lookups were affected, and later the feature-flag integration caused hanging requests and 504 responses. The lesson is that identity availability is an end-to-end property, not a single control outcome.

Practical implication: Model identity services as a dependency chain and test what happens when any upstream control plane is unavailable.

How cascading outages turn third-party identity dependencies into availability risk

The article shows a common failure mode in identity architecture: third-party services that are individually reasonable become collectively fragile when they share the same regional or platform dependency. AuthKit relied on a hosting provider that was also affected by the AWS event, and the feature-flag provider was still recovering when the second outage hit. That meant failover and recovery were not independent. Multi-region design only helps if the secondary path does not share the same hidden dependency pattern.

Practical implication: Map shared dependencies across hosting, flags, databases, and auth flows before assuming a secondary region is truly independent.

Why graceful degradation matters for hosted sign-in experiences

Hosted sign-in flows need explicit fallback behaviour because a blocked page load is not just a UX defect, it is an identity outage. The incident showed that sessions could remain valid while new sign-ins and page rendering failed, which is exactly the kind of split-state behaviour that resilience plans often miss. Timeouts, circuit breakers, and fallback logic are not embellishments here. They determine whether an identity product can continue to serve requests when one external service stops responding.

Practical implication: Design hosted authentication paths to degrade safely instead of waiting for every upstream dependency to recover.


NHI Mgmt Group analysis

Identity availability is now a control-plane governance problem, not just an infrastructure problem. This incident shows that authentication correctness does not preserve service continuity when credential lookup, hosting, and feature delivery all sit in the same failure domain. IAM teams should treat identity service resilience as part of the access architecture, because the user experience collapses before policy enforcement ever gets a chance to operate.

The assumption that a secondary region is an independent recovery path often breaks in practice. WorkOS attempted regional mitigation, but deployment failures and shared third-party dependencies limited its effectiveness. The hidden issue is not simply whether multi-region exists, but whether the failover path truly escapes the original dependency stack. Practitioners need to re-examine what they mean by redundancy when the same external services remain in the path.

Graceful degradation is the real availability boundary for modern identity services. A sign-in system can survive partial backend loss only if it can time out cleanly, fall back predictably, and continue core flows without waiting on optional integrations. The named concept here is identity dependency resilience: the ability to keep identity functions usable even when supporting services are impaired. That is the standard identity teams now need to engineer toward.

Session continuity is not the same as service availability. Existing sessions remained usable while new sign-ins and page loads failed, which can obscure the operational severity of an incident. For IAM programmes, this creates a misleading comfort gap: access may appear intact on the surface while the front door is effectively closed. Practitioners should distinguish between authenticated session persistence and real-time identity service uptime.

Identity service resilience now spans both IAM and SRE disciplines. The remediation set, including timeouts, circuit breakers, multiple regions, and graceful degradation, reflects an operational model where identity is treated as a business-critical service. That shift matters because governance has to account for service failure modes, not just access decisions. Teams that own identity platforms should align recovery objectives with the access experiences they are actually promising users.

What this signals

Identity service resilience: Identity programmes need to treat hosted sign-in, session handling, and dependent control-plane services as one operational surface. If the experience depends on a single region or a shared third-party path, the access layer is fragile even when the policy layer is sound.

The practical test is not whether authentication logic is correct, but whether the service can continue to serve requests when one provider degrades. That means validating failover, fallback, and timeout behaviour as part of identity governance, not as a separate reliability exercise.


For practitioners

  • Map identity dependency chains Inventory every upstream service that a sign-in, session refresh, or admin access flow depends on, including hosting, flags, database connectivity, and edge delivery. Mark which dependencies share the same failure domain so failover assumptions can be tested realistically.
  • Test graceful degradation paths Simulate unavailability of feature flags, database lookup, and hosted page rendering to confirm that the identity experience times out cleanly rather than hanging. Validate fallback behaviour for both login and existing session flows.
  • Separate session continuity from sign-in availability Track whether existing sessions continue, whether new authentication works, and whether management pages render normally during dependency outages. Treat those as distinct service outcomes, not one availability number.
  • Validate region independence end to end Review whether a secondary region actually avoids the same provider, deployment, or observability dependency that caused the original outage. A backup region that shares the same hidden control plane is not true resilience.
  • Set identity-specific recovery objectives Define recovery targets for sign-in, hosted authentication pages, and admin access separately from general infrastructure targets. Identity services need their own resiliency metrics because an outage here blocks access even when downstream systems remain healthy.

Key takeaways

  • Identity outages can block access even when authentication rules and sessions are functioning as designed.
  • The incident shows that hidden dependencies such as hosting, feature flags, and database lookup can turn one regional failure into a broader identity service disruption.
  • Teams should engineer graceful degradation and true multi-region independence into identity services, because resilience now defines whether IAM can actually deliver access.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0 and CIS Controls v8 set the governance and control requirements practitioners need to meet.

FrameworkControl / ReferenceRelevance
NIST CSF 2.0PR.AA-05 — Access Permissions, Entitlements and AuthorizationsIdentity service outages surfaced access path dependency on upstream authorization and credential retrieval.
RS.MI-01 — Incidents are containedThe article centers on containment and recovery after cascading service disruption.
PR.IR-01 — Network and Security Architectures Are ResilientThe core issue is whether identity architecture tolerates regional and provider failures.
Recommendation — Review access-path dependencies under PR.AA-05 and confirm auth flows still work when one control plane fails. Apply RS.MI-01 by validating containment and failover steps for identity outages before production incidents occur. Use PR.IR-01 to design identity services that remain available when one region or provider degrades.
CIS Controls v8CIS-13 — Network Monitoring and DefenseDegraded observability limited root-cause identification during the second outage.
Recommendation — Strengthen monitoring coverage so identity service degradation remains diagnosable when a provider is already impaired.

Key terms

  • Identity Service Resilience: Identity service resilience is the ability of sign-in, session, and access-control services to keep operating when a dependency fails. It includes fallback behaviour, failover design, and recovery mechanics that preserve access even when upstream systems are unavailable.
  • Graceful Degradation: Graceful degradation means a service continues to provide partial, predictable function when a dependency becomes unavailable. For identity systems, that might mean returning clean errors, preserving existing sessions, or falling back to cached state instead of hanging requests or breaking the login experience entirely.
  • Dependency Chain: A dependency chain is the set of direct and indirect libraries, runtimes, and services that an application relies on to operate. In practice, it determines whether an upgrade can happen cleanly or whether multiple coordinated changes are needed before the target version can be adopted safely.
  • Failover Path: The predefined route by which service traffic or workload processing shifts to a backup component after a failure. In messaging systems, a failover path is essential because it determines whether applications can continue operating when the primary broker becomes unavailable or degraded.

Deepen your knowledge

NHI governance, agentic AI identity, and machine identity lifecycle are core topics in our NHI Foundation Level course, the industry's only accredited NHI security programme. If you are responsible for identity security strategy or NHI governance in your organisation, it is worth exploring.
NHIMG Editorial Note
Published by the NHIMG editorial team on June 7, 2026.
Updated on October 7, 2026.
NHI Mgmt Group, the independent authority on Non-Human Identity, IAM, and Agentic AI security. nhimg.org