Cloud Concentration Risk: Why Cloud Outages Are Now Systemic

Diagram of four cloud risk layers (concentration, dependencies, control plane, data plane) with the control plane highlighted as the point of failure.
Cloud convergence has transformed the internet from a highly decentralised network into an ecosystem increasingly dominated by a handful of hyperscale providers. While this shift has enabled unprecedented scalability and innovation, it has also introduced systemic risks with far-reaching implications for critical infrastructure, artificial intelligence, and digital resilience. This article explores how cloud concentration is reshaping internet resilience and why modern outages are increasingly architecture failures rather than infrastructure failures.

Executive Summary

Over the past fifteen years, the services that run on the internet have consolidated onto a small number of hyperscale cloud platforms. That shift has brought scale, speed and a wave of innovation. It has also created a new kind of exposure: a fault in one provider’s control systems can now disrupt banks, airlines, hospitals and government services on several continents at once.

In this paper I look at how that concentration affects resilience, critical national infrastructure and emerging technologies such as AI. I set out a simple four-layer model for thinking about cloud systemic risk, test it against major outages from 2020 to 2026, and finish with the architectural and regulatory measures I believe organisations and policymakers should now prioritise.

Introduction

The internet was built to survive failure. Its designers wanted a network with no single point whose loss could bring the whole thing down. Traffic could be rerouted, networks could heal around damage, and services could stay available even when parts of the infrastructure did not. For decades, that design underpinned a remarkably dependable digital economy.
Something has shifted. The network itself is still broadly decentralised, but the platforms, applications and utilities that sit on top of it have become concentrated in the hands of a few cloud providers. Payment systems, consumer apps, public services and, increasingly, AI workloads all depend on the same small group of hyperscalers, and a large share of global digital activity now passes through their infrastructure.
On most measures this is a success story. Cloud computing has made technology cheaper, quicker to deploy and easier to scale, and it has made entirely new categories of service possible. It has also produced a paradox that I think is under-appreciated: the internet remains distributed in design, yet it is increasingly centralised in practice.
In my view, the biggest risk is not the cloud as such. It is the centralisation hidden inside cloud architectures, particularly in the control layers that decide how systems find, authenticate and communicate with each other. When those layers fail, the damage does not stay local.
Recent years have shown this repeatedly. Faults in apparently minor components (a DNS record, an identity service, a routing configuration) have cascaded across regions and industries, disrupting everything from card payments to emergency services. I do not see these as isolated accidents. They point to a structural problem, and understanding it means looking beyond where systems run to how they are controlled.
The rest of this paper traces how we arrived here, offers a practical framework for thinking about cloud systemic risk, and argues that resilience planning must now account for failure domains defined by architecture rather than geography.

Origins of the Internet

No single organisation owns the internet. The name comes from “internetworking”: it is a mesh of tens of thousands of independently operated networks that agree to exchange traffic using common protocols. That is why no company, country or region can lay claim to it, and why it has become the backbone for almost every business and public service.
Decentralisation ran through both its architecture and its governance. Routes, paths and naming services were spread across many operators and locations, so that a failure in one place could be worked around.
The commercial internet grew out of ARPANET and the academic research networks. As these opened to commercial traffic in the early 1990s, independent internet service providers (ISPs) began selling access to individuals, businesses and organisations over dial-up, DSL, cable, satellite and, later, fibre.

Each ISP ran its own autonomous system (AS). Early routing relied on routes configured by hand between known hosts. This gave way to dynamic routing, with Interior Gateway Protocols (IGPs) handling routes inside a provider’s network and the Border Gateway Protocol (BGP) handling routes between providers. ISPs are usually described in three tiers:

  • Tier 1: global backbone providers that exchange traffic with each other without paying for transit.
  • Tier 2: regional providers that buy transit from Tier 1 networks and peer locally.
  • Tier 3: local ISPs serving end users.

    In the early 2000s, many telecommunications companies (telcos) moved into the ISP market, and some sold parts of their voice business to fund investment in IP networks. Between them, ISPs and telcos ran the large data centres that hosted websites, enterprise applications and connectivity services.
    Internet exchange points (IXPs) are the physical locations where networks meet to exchange traffic directly, while points of presence (PoPs) are where a provider’s network connects to other networks or to customers. Both are spread across major cities around the world.

The Rise of Cloud

The terms “the cloud” and “the internet” are often used interchangeably. They are not the same thing, and the difference matters for the argument that follows.
The idea behind cloud computing predates the industry. In 1961 John McCarthy suggested that computing might one day be sold as a public utility, and the phrase “cloud computing” appears in a Compaq strategy document from 1996, borrowed from the cloud symbol engineers used to represent the network in diagrams [1].
The most widely used definition comes from the US National Institute of Standards and Technology (NIST), which identifies five essential characteristics [2]:

  • On-demand self-service: customers provision computing, storage and network resources themselves, without needing to deal with anyone at the provider.
  • Broad network access: services are reached over the network through standard mechanisms, from phones, laptops, workstations and other systems.
  • Resource pooling: the provider’s physical and virtual resources serve many customers through a multi-tenant model and are reassigned according to demand. Customers generally do not control exactly where resources sit, although they may be able to choose a region or country.
  • Rapid elasticity: capacity can be scaled up or down quickly, often automatically, so that to the customer it can appear almost unlimited.

Measured service: usage is metered, so it can be monitored, controlled and reported to both provider and customer.
Cloud computing in its modern form, meaning servers, storage, databases, networking, software and analytics delivered over the internet, took shape in the 2000s. Instead of owning hardware and data centres, organisations rent what they need and pay for what they use.

2002: Amazon releases its first web services for developers.

2006: AWS launches S3 and EC2, widely regarded as the start of the modern public cloud.

2008–2010: Google App Engine and Microsoft Azure launch.

Underneath it all are data centres in which compute, storage and networking are pooled and virtualised so they can be sold in small slices. Services are usually grouped into compute, storage, networking and higher-level application services.

Today the cloud and the internet are so intertwined that most users cannot tell them apart. The important difference is ownership. The internet has no owner, whereas the cloud platforms that now carry so much of its traffic are owned and run by a handful of private companies.

Modern Internet Cloud

Legacy internet was autonomous and decentralised without political or commercial ownership or control whereas the cloud is owned and managed by a handful of private business entities. 

What’s Changed

Traditional Internet

Modern Cloud Internet

Thousands of ISPs

Few hyperscalers

Thousands of hosting providers

Shared platforms

Many DNS providers

Shared identity systems

Distributed ownership

Shared control planes

To be precise, the internet itself has not become centralised; the services running on it have. Cloud providers may rise and fall, but the underlying network will outlast them.

Cloud Convergence and Centralisation

In the early commercial internet, hosting and data centre services were supplied by thousands of ISPs and hosting companies spread across the world. Many of those firms have since been consolidated, acquired or out-competed, and most new workloads now land on a small number of hyperscale providers.
The benefits are real, but so is the cost: concentration has created single points of failure on a global scale. The hyperscalers now run their own private global backbones, edge PoPs, DNS services and content delivery networks, and connect directly with other networks at exchange points. A growing share of traffic therefore travels over infrastructure operated end to end by one company, even though telcos still provide most last-mile access.
The market data makes the point. Synergy Research Group estimates that in the first quarter of 2026 Amazon held 28% of the worldwide cloud infrastructure services market, Microsoft 21% and Google 14%. Together that is 63% of a market worth almost US$129 billion in that quarter alone [3]. No other provider held more than 4%.

Worldwide market share of leading cloud infrastructure service providers in Q1 2026

Cloud Hyper Scalers

Source: Synergy Research Group, chart via Statista [4].

Cloud is no longer optional; it is foundational. Gartner forecast worldwide end-user spending on public cloud services of US$723.4 billion in 2025, and predicts that 90% of organisations will adopt a hybrid cloud approach through 2027 [5].
Because of that concentration, cloud platforms now underpin sectors that society cannot do without, including:

  • Financial services and payments
  • Healthcare
  • Government and public services
  • Communications
  • AI and data processing.

    A major cloud outage is therefore no longer just a technical incident. It is a systemic one. When a single provider falters, the effects can reach hospitals, transport networks, energy and water utilities, emergency services and financial markets, often simultaneously. The risk extends beyond commercial losses to national infrastructure, public trust and, in the worst cases, human safety.
    This paper is about those events and their growing frequency. The question is no longer whether major outages will happen, but how badly they will affect a world that increasingly runs on the cloud.
    Infrastructure has become more reliable, but outages remain common and expensive. Uptime Institute’s 2026 analysis found that [6]:

  • 57% of operators said their most recent major outage cost more than US$100,000, and one in five put the cost above US$1 million
  • Over the nine years Uptime has tracked publicly reported outages, third-party IT and data centre providers (including cloud and internet companies, telcos and colocation firms) have accounted for about two-thirds of them
  • Failure to follow established procedures remains the leading driver of outages caused by human error.
    Because modern cloud systems are so interconnected, failures rarely stay contained. They spread across services, regions and industries.

    The central argument of this paper is that major cloud failures are increasingly failures of architecture rather than of infrastructure.

The Cloud Systemic Risk Model

Cloud concentration risk is often treated as a simple availability question: will the servers stay up? In my experience the real risk comes from the interaction of four layers: the data plane, the control plane, the dependency layer and the concentration layer.

Diagram of four cloud risk layers (concentration, dependencies, control plane, data plane) with the control plane highlighted as the point of failure.The key point of the model is that the most damaging cloud failures now tend to start in the control plane rather than in the underlying hardware. As dependencies accumulate and ownership concentrates, a fault in one of these layers can travel well beyond where it started.
Data plane

Where applications, workloads and data actually run: virtual machines, databases, containers and storage. Failures here tend to be local.

Control plane
DNS, identity and access management, routing, API gateways and orchestration. A failure here can bring down systems that are otherwise perfectly healthy, which is why I regard this as the layer carrying the greatest systemic risk.

Dependency layer
Very few services stand alone. A typical application is built from microservices, APIs and SaaS integrations, and may depend at the same time on Azure, Microsoft 365, Salesforce, Stripe, Cloudflare and Okta. Failures travel along those dependencies, often by routes the application owner never mapped.

Concentration layer
The ownership and control of infrastructure by a small number of hyperscalers. This is the strategic layer. Even where individual services are well engineered, AWS, Azure and Google Cloud between them host a very large share of the world’s workloads, so their failures quickly become everyone’s failures.

Outage Analysis: 2020 to 2026

The incidents below are drawn from the providers’ own post-incident reports wherever these were published. Taken together, they show how often the trigger is a configuration change or a control-plane fault rather than a hardware failure.

2026

Azure virtual machines and managed identities (2–3 February 2026)

A storage access policy intended to block anonymous access was wrongly applied to storage accounts the platform itself relied on. Virtual machines could no longer retrieve the extension files they needed, and the problem cascaded into failed VM and Kubernetes deployments, managed identity errors in US regions and broken CI/CD pipelines, lasting about 12 hours [7].

  • Root cause: policy change applied to the wrong resources
  • Layer affected: control plane

2025

AWS US-EAST-1 DynamoDB DNS failure (19–20 October 2025)

A race condition in DynamoDB’s automated DNS management left the service’s regional endpoint with an empty DNS record that the automation could not repair. Anything that needed to reach DynamoDB in Northern Virginia failed, and the fault spread to EC2 instance launches, Network Load Balancer, Lambda, ECS, EKS, Amazon Connect, STS and other services. Full recovery took more than 14 hours [8].

  • Root cause: defect in DNS automation (control plane)
  • Impact: AWS services and a long list of customer applications across finance, media, retail and enterprise IT

What makes this incident significant is that the servers were largely fine. The outage came from a control-plane dependency in a single region, which rippled out to services around the world that relied on it. It is a clear case of an orchestration failure rather than an infrastructure failure.

Azure Front Door (29–30 October 2025)

Nine days later, Microsoft’s global edge and content delivery service went down for more than eight hours. A sequence of customer configuration changes, processed by two different versions of the control plane, produced incompatible metadata that exposed a latent bug and crashed the data plane. The Azure Portal, Entra ID, Microsoft 365, Azure SQL Database, App Service and many customer websites were affected [9].

  • Root cause: configuration change exposing a software defect
  • Impact: Microsoft 365, Azure services and customer websites worldwide

The lesson is that global edge and entry-point services are critical failure domains in their own right.

Google Cloud Service Control (12 June 2025)

A new quota policy feature in Service Control, the system that checks API requests, lacked proper error handling and was not protected by a feature flag. When a policy update containing blank fields was replicated globally, it triggered a null pointer exception that crashed Service Control in every region. Dozens of Google Cloud and Google Workspace products returned errors for several hours [10].

  • Root cause: policy change exposing an untested code path
  • Impact: API failures across Google Cloud and Workspace products worldwide.

Cloudflare 1.1.1.1 resolver (14 July 2025)

A configuration error made in June had wrongly linked the IP prefixes of Cloudflare’s public DNS resolver to a pre-production service. When that service was reconfigured on 14 July, the resolver’s routes were withdrawn globally and 1.1.1.1 was unreachable for 62 minutes. Cloudflare confirmed that this was an internal error, not an attack [11].

  • Root cause: configuration error
  • Impact: DNS resolution failures for 1.1.1.1 users worldwide

Cloudflare global network (18 November 2025)

A database permissions change caused a Bot Management configuration file to double in size. When the oversized file was pushed across Cloudflare’s network, it exceeded a hard limit in the proxy software, which then failed and returned errors for a large proportion of traffic over roughly six hours [12]. Because so many websites and applications sit behind Cloudflare, the impact was felt well beyond its direct customers.

  • Root cause: faulty configuration file generation
  • Impact: widespread website and application failures worldwide

2024.

CrowdStrike Falcon update (19 July 2024)

Strictly speaking this was not a cloud outage, but it is one of the clearest demonstrations of shared dependency. A faulty content update to CrowdStrike’s Falcon sensor crashed about 8.5 million Windows devices, which Microsoft estimated at less than 1% of all Windows machines [13]. Airlines, hospitals, broadcasters and banks were among those hit. Parametrix estimated the direct losses to US Fortune 500 companies, excluding Microsoft, at US$5.4 billion [14].

  • Root cause: faulty software update.
  • Impact: airlines, healthcare, financial services and media

Google Cloud, Frankfurt (24 October 2024)

A power failure and cooling problem took part of the europe-west3-c zone offline for around 12 hours, cutting European customers off from their virtual machines and disks [15].

The big three across 2024
Parametrix found that critical service interruptions across AWS, Azure and Google Cloud rose by 18% in 2024, and recorded six events that each lasted more than ten hours, with nearly 100 hours of downtime between them [16].

2023

  • Microsoft Azure WAN (25 January 2023):

    an engineer ran a command intended for a single router that instead triggered a recomputation across every router in Microsoft’s wide-area network. Users of Azure, Microsoft 365 and Power Platform around the world saw timeouts and connectivity failures for several hours [17].

  • Putting a price on concentration:

    Parametrix modelled a 24-hour outage of critical services in AWS US-EAST-1 at US$3.4 billion in direct revenue losses for Fortune 500 companies, rising to US$7.8 billion over 48 hours [18].

2022

  • Google Cloud, London (19 July 2022):

    during the UK’s record heatwave, several redundant cooling systems failed at the same time in a europe-west2 data centre. Google had to power down equipment, and a routing error during the response made regional storage services unavailable across all three zones [19].

  • AWS US-EAST-2 (28 July 2022):

    a power loss in a single availability zone disrupted 38 AWS services, including EC2 and DynamoDB, for nearly three hours [20].

2021

  • Microsoft Azure DNS (1 April 2021):

    an unusual global surge in DNS queries, combined with a bug in the code designed to absorb such surges, overwhelmed Azure DNS for about 40 minutes and disrupted Microsoft and customer services [21].

  • AWS US-EAST-1 (7 December 2021):

    an automated scaling activity triggered a flood of connections that overwhelmed networking devices between AWS’s internal and main networks, disrupting AWS services and a wide range of consumer platforms for several hours [22].

2020

  • AWS Kinesis, US-EAST-1 (25 November 2020):

    adding capacity to the Kinesis front-end fleet pushed every server past an operating system thread limit. Kinesis failed and took dependent AWS services and customer applications with it for several hours [23].

  • Google authentication (14 December 2020):

    an error in a quota system during a migration left Google’s User ID Service unable to write data. For about an hour, services that required users to sign in, including Gmail, YouTube, Workspace and parts of Google Cloud, returned errors [24].

Implications for Critical Sectors

Healthcare
Outages can delay access to patient records, disrupt scheduling systems and slow emergency care. The CrowdStrike incident showed how quickly clinical services are forced back onto paper when shared platforms fail.

Finance and the wider economy
Banking platforms, payment systems and trading venues depend on real-time data. Downtime can freeze transactions and halt commerce: businesses lose revenue, supply chains stall and customers lose access to services.

Government, Defence and Social order
Because identity systems and public services are increasingly cloud-hosted, outages now carry national security and governance implications. Military communications, intelligence analysis and logistics also make growing use of commercial cloud platforms. This raises several concerns:

  • Strategic vulnerability: an outage could impede command and control, intelligence sharing or the operation of autonomous defence systems.
  • National security: hostile states may see cloud infrastructure as an attractive target for disrupting national capabilities.
  • Public trust: if public services and emergency response falter during an outage, confidence in institutions suffers, and that matters most during a crisis.
  • Communication: when messaging, collaboration and social platforms go dark, personal and professional communication goes with them.

AI systems
AI training and inference are increasingly concentrated in cloud-hosted GPU clusters. Autonomous systems, from automated trading engines to smart-city management and vehicle fleets, rely on the cloud for computation, data and coordination. An outage or serious degradation can lead to:

  • Operational paralysis: without access to data or processing power, AI-driven systems may stop without warning, causing safety incidents or financial losses.
  • Data integrity problems: interrupted synchronisation or corrupted data can undermine the performance and trustworthiness of models.
  • Security gaps: failing over to less secure backup systems during an outage can open doors for attackers.

Because so many of these systems rely on the same few providers, a single disruption can affect all of them at once.

Critical national infrastructure

  • Energy: cloud-based control and analytics for electricity distribution and smart grids depend on real-time data. A major outage could interrupt monitoring, coordination and blackout response.
  • Transport and logistics: airlines, railways and supply chains rely on cloud-managed scheduling and tracking, so an outage can quickly turn into widespread delays and economic losses.
    In each case the pattern is the same: one point of failure, many real-world consequences.

Root Causes of Modern Cloud Failure

The case studies above share a striking feature: very few began with hardware failing on its own. Most started with software, configuration or operational processes inside highly complex distributed systems. The common triggers are set out below.

Configuration errors and human mistakes
The Azure WAN outage, Azure Front Door, Cloudflare’s resolver incident and Google’s Service Control failure all began with a change that behaved differently from what its authors expected [9, 10, 11, 17]. Uptime Institute reports that nearly 40% of organisations have suffered a major outage caused by human error in the past three years, and that 85% of those incidents stem from staff not following procedures or from flaws in the procedures themselves [25].

Software defects and faulty updates
A faulty update or latent bug pushed through a global deployment pipeline can spread worldwide within minutes, as the CrowdStrike incident and Cloudflare’s November 2025 outage both showed [12, 13].

Network and control-plane failures
Routing errors, DNS faults and broken identity services can cut off whole regions even when the compute underneath is healthy. Uptime found that IT and networking issues rose in 2024 to account for 23% of impactful outages [25].

Complexity, capacity and interdependence
Unexpected surges in demand, legitimate or malicious, can overwhelm capacity; the 2021 Azure DNS and AWS network incidents both began this way [21, 22]. And because services depend on each other in ways that are rarely fully mapped, a fault in one component can spread quickly.

Emerging Systemic Risks

Taken together, cloud convergence creates a distinct category of risk made up of:

  • Dependence on a single provider
  • Control-plane centralisation that customers cannot see
  • Failure domains that are global rather than regional
  • Chains of service dependencies that cross provider boundaries.

Uptime Institute’s data points the same way: third-party providers, including the large cloud and internet companies, account for about two-thirds of publicly reported outages [6]. This complexity makes failures harder to predict, contain and recover from.

Future Risk Outlook

Looking ahead, I see four trends that are likely to increase systemic cloud risk over the next decade.

AI centralisation
Training and inference workloads are concentrating in a small number of hyperscale GPU clusters.

Deeper dependency in critical sectors
Healthcare, energy and transport continue to move core systems onto cloud platforms.

Growing interconnection
Modern applications depend on hundreds of interconnected services, each one a potential path for an outage to travel.

Sovereignty and geopolitics
Many countries depend on foreign-owned cloud infrastructure, which raises strategic and regulatory questions that are only starting to be addressed.

Together, these trends mean that future outages could have larger social and economic consequences than any we have seen so far.

Mitigation Strategies

My recommendation is that organisations adopt resilience-first architectures built on six principles.

  1. Multi-cloud where it counts
    Spread critical workloads across providers to reduce dependence on any one of them, starting with the services whose loss you could not tolerate.
  2. Active-active architectures
    Avoid relying on cold failover. Run critical systems across regions, and where justified across providers, at the same time.
  3. Decoupled identity and DNS
    Reduce reliance on a single provider’s control-plane services, particularly for authentication and name resolution.
  4. Chaos engineering and resilience testing
    Test failure scenarios regularly, including the loss of a provider’s control plane, rather than assuming that failover will work when it is needed.
  5. Regulatory oversight
    Treat major cloud platforms as critical infrastructure subject to resilience standards. The UK’s critical third parties regime for the financial sector and the EU’s Digital Operational Resilience Act (DORA) are early steps in this direction [26, 27].
  6. Edge computing
    Move processing closer to users and devices so that essential functions can keep running when central services are unavailable.

What Most Organisations Get Wrong About Cloud Resilience

Many organisations assume that spreading workloads across multiple availability zones makes them resilient. In practice, those zones often share control planes, identity systems, DNS and operational processes. The October 2025 AWS incident is a good example [8]: redundancy within a single provider does not remove concentration risk.

Limitations

My analysis relies mainly on publicly reported outages, providers’ own incident reports and industry research. The true frequency, causes and cost of cloud failures are probably under-reported, because providers do not publish full details of every event. It is also worth saying plainly that hyperscale cloud remains more reliable than many traditional enterprise environments. This paper is about systemic concentration risk, not about whether the cloud is reliable overall.

Conclusion

Cloud computing has transformed the internet and made innovation possible at a scale that would have been hard to imagine twenty years ago. That transformation has come with a trade-off: centralisation has increased systemic risk.
Technology leaders, policymakers and risk managers need to work together to diversify digital supply chains, set and enforce resilience standards, and make sure the infrastructure we are building our future on can withstand failure.
Efficiency and resilience have to be balanced. Without deliberate architectural, organisational and regulatory action, the forces that made the cloud revolution possible could become its greatest weakness.

Frequently Asked Questions

What is cloud concentration risk?

Cloud concentration risk is the danger that comes from so many organisations depending on the same few cloud providers. In Q1 2026, AWS, Microsoft and Google held 63% of the worldwide cloud infrastructure market between them. A single fault at one provider can therefore disrupt thousands of businesses and public services at once.

Why do cloud outages now have a global impact?

Most major outages now start in the control plane: the DNS, identity, routing and orchestration systems that tell services how to find and trust each other. Many services worldwide share these systems. A fault in one region can therefore spread to applications that were otherwise healthy, across industries and continents.

What caused the AWS outage in October 2025?

On 19–20 October 2025, a race condition in DynamoDB’s automated DNS management left its US-EAST-1 endpoint with an empty DNS record. Services that depended on DynamoDB failed in turn, including EC2 launches, Lambda and Network Load Balancer. Full recovery took more than 14 hours, even though the underlying servers were largely unaffected.

Does deploying across multiple availability zones make my systems resilient?

Not on its own. Availability zones protect against local hardware, power and cooling failures. However, zones within a provider usually share the same control plane, identity services and DNS. If one of those fails, every zone can be affected at once. Redundancy within a single provider does not remove concentration risk.

Is a multi-cloud strategy worth the cost?

For your most critical services, often yes. Running everything across several providers is expensive and complex. A targeted approach is usually better value: identify the services whose loss you could not tolerate, then run those active-active across regions or providers. Keep identity and DNS independent of any single provider’s control plane.

How are regulators responding to cloud concentration risk?

The UK has introduced a critical third parties regime, under which HM Treasury can designate major technology providers to the financial sector for direct oversight by the Bank of England, the PRA (Prudential Regulation Authority) and the FCA. In the EU, the Digital Operational Resilience Act (DORA) has applied since January 2025. It includes oversight of critical ICT (information and communications technology) third-party providers. Both are early steps towards treating cloud platforms as critical infrastructure.

Is the cloud less reliable than running my own data centre?

Generally, no. Hyperscale cloud is more reliable than many traditional enterprise environments. The concern is not how often each provider fails. It is how widely a failure spreads when so many organisations depend on the same platforms. That is a systemic risk, not a reliability problem.

Where should an organisation start?

Start by mapping your dependencies: which cloud regions, identity providers, DNS services and SaaS platforms your critical services rely on, including indirect ones. Then test what actually happens when each one fails. Most organisations find at least one hidden single point of failure they had not planned for.

References

[1] Regalado, A. (2011) ‘Who coined “cloud computing”?’, MIT Technology Review, 31 October. Available at: https://www.technologyreview.com/2011/10/31/257406/who-coined-cloud-computing/ (accessed 29 September 2026).

[2] Mell, P. and Grance, T. (2011) The NIST Definition of Cloud Computing, NIST Special Publication 800-145. Gaithersburg, MD: National Institute of Standards and Technology. Available at: https://doi.org/10.6028/NIST.SP.800-145 (accessed 29 September 2026).

[3] Synergy Research Group (2026) ‘Cloud market annual revenue run rate topped half a trillion dollars in Q1 as growth surge continues’, 29 April. Available at: https://www.srgresearch.com/articles/cloud-market-annual-revenue-run-rate-topped-half-a-trillion-dollars-in-q1-as-growth-surge-continues (accessed 29 September 2026).

[4] Statista (2026) ‘Big Three hold dominant lead in accelerating cloud market’ [chart], based on Synergy Research Group data for Q1 2026. Available at: https://www.statista.com/chart/18819/worldwide-market-share-of-leading-cloud-infrastructure-service-providers/ (accessed 29 September 2026).

[5] Gartner (2024) ‘Gartner forecasts worldwide public cloud end-user spending to total $723 billion in 2025’, press release, 19 November. Available at: https://www.gartner.com/en/newsroom/press-releases/2024-11-19-gartner-forecasts-worldwide-public-cloud-end-user-spending-to-total-723-billion-dollars-in-2025 (accessed 29 September 2026).

[6] Uptime Institute (2026) ‘Uptime announces Annual Outage Analysis Report 2026’, press release, 13 May. Available at: https://uptimeinstitute.com/about-ui/press-releases/uptime-announces-annual-outage-analysis-report-2026 (accessed 29 September 2026).

[7] Microsoft Azure (2026) ‘Post Incident Review: Virtual Machines, Managed Identities, and dependent services – multi-region outage’, Azure status history, tracking ID FNJ8-VQZ. Available at: https://azure.status.microsoft/en-us/status/history/?trackingId=FNJ8-VQZ (accessed 29 September 2026).

[8] Amazon Web Services (2025) ‘Summary of the Amazon DynamoDB service disruption in the Northern Virginia (US-EAST-1) Region’, October. Available at: https://aws.amazon.com/message/101925 (accessed 29 September 2026).

[9] Microsoft Azure (2025) ‘Post Incident Review: Azure Front Door – connectivity issues across multiple regions’, Azure status history, tracking ID YKYN-BWZ. Available at: https://azure.status.microsoft/en-us/status/history/?trackingId=YKYN-BWZ (accessed 29 September 2026).

[10] Google Cloud (2025) ‘Multiple GCP products service disruption’, incident report, 12 June. Available at: https://status.cloud.google.com/incidents/ow5i3PPK96RduMcb1SsW (accessed 29 September 2026).

[11] Pallarito, A. and Abley, J. (2025) ‘Cloudflare 1.1.1.1 incident on July 14, 2025’, Cloudflare Blog, 15 July. Available at: https://blog.cloudflare.com/cloudflare-1-1-1-1-incident-on-july-14-2025 (accessed 29 September 2026).

[12] Prince, M. (2025) ‘Cloudflare outage on November 18, 2025’, Cloudflare Blog, 18 November. Available at: https://blog.cloudflare.com/18-november-2025-outage/ (accessed 29 September 2026).

[13] Microsoft (2024) ‘Helping our customers through the CrowdStrike outage’, The Official Microsoft Blog, 20 July. Available at: https://blogs.microsoft.com/blog/2024/07/20/helping-our-customers-through-the-crowdstrike-outage/ (accessed 29 September 2026).

[14] Parametrix (2024) ‘CrowdStrike to cost Fortune 500 $5.4 billion; insured loss range of $540 million to $1.08 billion’, July. Available at: https://www.parametrixinsurance.com/in-the-news/crowdstrike-to-cost-fortune-500-5-4-billion-insured-loss-range-of-540-million-to-1-08-billion (accessed 29 September 2026).

[15] DatacenterDynamics (2024) ‘Google Cloud region in Germany suffers 12-hour outage’, 25 October. Available at: https://www.datacenterdynamics.com/en/news/google-cloud-region-in-germany-suffers-12-hour-outage/ (accessed 29 September 2026).

[16] Parametrix (2025) ‘“Critical” cloud service outages up by nearly one fifth in 2024’. Available at: https://www.parametrixinsurance.com/in-the-news/2024-cloud-outage-risk-report (accessed 29 September 2026).

[17] Microsoft Azure (2023) ‘Post Incident Review: Azure Networking – global WAN issues’, Azure status history, tracking ID VSG1-B90. Available at: https://azure.status.microsoft/en-us/status/history/?trackingId=VSG1-B90 (accessed 29 September 2026).

[18] Parametrix (2023) ‘Fortune 500 companies vulnerable to cloud risks, could suffer losses of $20 billion’, 17 November. Available at: https://www.parametrixinsurance.com/in-the-news/fortune-500-vulnerable-to-cloud-risks-could-suffer-losses-of-20-billion (accessed 29 September 2026).

[19] Swinhoe, D. (2022) ‘Google’s London data center outage during heatwave caused by “simultaneous failure of multiple, redundant cooling systems”’, DatacenterDynamics, 2 August. Available at: https://www.datacenterdynamics.com/en/news/googles-london-data-center-outage-during-heatwave-caused-by-simultaneous-failure-of-multiple-redundant-cooling-systems/ (accessed 29 September 2026).

[20] Data Center Knowledge (2023) ‘Update: AWS experienced an outage in its US-East 2 availability zone’, 25 January. Available at: https://www.datacenterknowledge.com/outages/update-aws-experienced-an-outage-in-its-us-east-2-availability-zone (accessed 29 September 2026).

[21] Redmond, T. (2021) ‘Azure’s DNS problem highlights fallibility in cloud services’, Practical365, 9 April. Available at: https://practical365.com/azure-dns-failure-highlights-cloud-failibility/ (accessed 29 September 2026).

[22] Amazon Web Services (2021) ‘Summary of the AWS service event in the Northern Virginia (US-EAST-1) Region’, December. Available at: https://aws.amazon.com/message/12721 (accessed 29 September 2026).

[23] Amazon Web Services (2020) ‘Summary of the Amazon Kinesis event in the Northern Virginia (US-EAST-1) Region’, November. Available at: https://aws.amazon.com/message/11201 (accessed 29 September 2026).

[24] Wheatley, M. (2020) ‘Google blames last week’s outage on Google User ID Service error’, SiliconANGLE, 23 December. Available at: https://siliconangle.com/2020/12/23/google-blames-last-weeks-outage-google-user-id-service-error/ (accessed 29 September 2026).

[25] Uptime Institute (2025) ‘Uptime announces Annual Outage Analysis Report 2025’, press release, 6 May. Available at: https://uptimeinstitute.com/about-ui/press-releases/uptime-announces-annual-outage-analysis-report-2025 (accessed 29 September 2026).

[26] Bank of England, Prudential Regulation Authority and Financial Conduct Authority (2024) PS16/24: Operational resilience – critical third parties to the UK financial sector, November. Available at: https://www.bankofengland.co.uk/prudential-regulation/publication/2024/november/operational-resilience-critical-third-parties-to-the-uk-financial-sector-policy-statement (accessed 29 September 2026).

[27] European Union (2022) Regulation (EU) 2022/2554 on digital operational resilience for the financial sector (DORA), Official Journal of the European Union, L 333, 27 December. Available at: https://eur-lex.europa.eu/eli/reg/2022/2554/oj (accessed 29 September 2026).

#CloudComputing #DigitalInfrastructure #CyberSecurity #TechStrategy #EnterpriseIT

Related Posts

Join Our Newsletter