Join our Newsletter — 33% off our NHI Course

Why does concentration risk matter for resilience planning?

Concentration risk matters because a small number of providers can create a large shared-fate surface. When many organisations build on the same infrastructure, one outage can affect multiple tenants, suppliers, and customers at once. Resilience planning has to account for that shared exposure, not just for single-system uptime.

Why concentration risk changes the resilience equation

concentration risk is not just a scale issue, it changes the failure model. If many workloads, suppliers, or customers depend on the same provider, region, platform, or control plane, then a single fault can become a correlated outage. Resilience planning has to measure shared-fate exposure, not assume that repeated use of the same service is “diversified” by volume.

The practical implication is that availability assumptions built around one system often fail when the dependency is systemic. A narrow supplier base can amplify blast radius, reduce failover options, and turn a routine disruption into an ecosystem event. The question is not whether any one platform is reliable, but how much of the business collapses when that platform degrades.

That is why concentration risk belongs in resilience design, architecture review, and third-party oversight together. It affects recovery objectives, substitute capacity, and the order in which dependencies can be restored when multiple parties are affected at the same time.

What resilience teams should assess before they trust a dependency

Resilience planning should start by mapping where correlated failure is possible. The most useful view is often not “which systems are critical?” but “which systems fail together?” That includes common cloud regions, shared identity providers, managed network paths, shared software dependencies, and single providers supporting multiple business functions.

Testing should also reflect the dependency, not just the application. If failover lands on the same provider, same region, or same control plane, then the test proves continuity only on paper. Good resilience work checks for independence of recovery paths, capacity to absorb migration, and whether alternative processing can actually run without the shared service.

For organisations in regulated sectors, the issue is formalised in third-party and operational resilience expectations. The EU Digital Operational Resilience Act (DORA) places ICT third-party risk and resilience testing at the centre of planning, while the NIST Cybersecurity Framework 2.0 frames recovery as part of an end-to-end resilience lifecycle.

How concentration risk shows up in incident response and recovery

Concentration risk makes incidents harder to contain because the outage boundary is wider than the local environment. When the same provider supports many tenants or business units, an incident can affect onboarding, authentication, data access, vendor operations, and customer-facing services at once. Recovery may also be constrained by vendor prioritisation, throttling, or delayed restoration of shared dependencies.

This is why backup design alone is insufficient. A backup is only resilient if the restore path is independent enough to work during the same event that took the primary path down. Teams should also expect that a shared provider event can create concurrency problems, where many organisations compete for support, compute, or network capacity at the same time.

Operationally, the response plan should identify which functions can be degraded, which must be isolated first, and which business processes need manual workarounds while the shared dependency recovers. The right planning unit is not just a system, but the set of downstream services that move together when the common provider fails.

Risk and Threat Considerations

Concentration risk creates a shared-fate failure mode, so a single outage, misconfiguration, or provider-level disruption can propagate across many organisations at once. That raises the likelihood of simultaneous service loss, slower restoration, and broader business interruption than a localised failure would create.

Failure mechanism: Correlated dependency on one provider, region, or platform removes independence from the recovery path, so failover capacity, incident response, and restore options collapse together when the shared layer degrades.

Impact: Organisations can lose availability across multiple services at once, miss recovery targets, and discover too late that their continuity design did not actually provide an independent alternative.

Standards & Framework Alignment

This section maps relevant standards and security frameworks to the operational risks and controls described in this guidance.

NIST CSF 2.0 and CIS Controls v8 set the technical controls, while DORA defines the regulatory obligations.

Framework Control / Reference Relevance
NIST CSF 2.0 RC.RP-01 — Recovery Plan Execution Concentration risk changes how recovery must be executed under shared-fate outages.
GV.SC-01 — Cyber Supply Chain Risk Management Strategy Shared provider dependence is a supply-chain resilience problem.
Recommendation — Design recovery paths that remain usable when the shared dependency is the outage source. Map concentration exposure across suppliers and critical services, then reduce correlated dependency.
DORA ICT third-party risk and operational resilience DORA directly addresses shared provider risk, resilience testing, and third-party dependency exposure.
Recommendation — Test whether third-party concentration can break continuity and document substitute arrangements.
CIS Controls v8 CIS-17 — Incident Response Management Shared-fate events change the scope and coordination needs of incident response.
Recommendation — Prepare incident playbooks for provider-wide outages that affect multiple business services at once.

Practitioner Guidance

What to prioritise: Rank dependencies by shared-fate exposure, not by vendor name alone. The highest-risk items are the ones whose failure would simultaneously affect multiple critical services, not the ones with the largest contract value.

What to verify: Confirm that recovery paths are genuinely independent for network, identity, storage, and control-plane dependencies. If the backup relies on the same provider class or region, treat it as a correlated dependency, not a separate control.

What good looks like: A resilient design can show that one provider failure does not automatically take down every important workload, and that degraded operations can continue long enough for recovery to complete.

Practitioner takeaway: Concentration risk matters because resilience is only real when failure domains are independent enough to give you an actual second chance, not just a second label for the same dependency.