Executive Summary
Cloud convergence has transformed and reshaped the Internet from a highly decentralised system into a concentrated ecosystem dominated by a small number of hyperscale operators. While this shift has delivered scalability, performance and innovation, it has also introduced systemic risks where outages can cascade globally.
In this paper, I examine how cloud centralisation impacts resilience, critical infrastructure as well as emerging technologies such as AI and proposes multi-cloud and regulatory strategies to mitigate these risks.
Introduction
In its original architecture, the internet was built to survive failure. It was decentralised, resilient and adaptive, designed to ensure that no single disruption could bring down the system as a whole; because data could be rerouted, networks could self-correct and services could remain available even when parts of the infrastructure failed. For decades, this model underpinned the growth of a robust and reliable digital ecosystem. But something fundamental has changed.
Over the last decade, the centre of gravity has shifted from distributed networks to centralised services. While the underlying network topology remains largely decentralised, the utilities, applications, platforms and systems that power modern digital life have become increasingly concentrated within a small number of cloud providers.
This shift has been subtle but profound. Today, most business-critical systems, consumer platforms and emerging technologies from financial systems to AI workloads depend on hyperscale cloud infrastructure. What was once a diverse and distributed ecosystem has evolved into one where a significant proportion of global digital activity is mediated through a handful of platforms.
At first glance, this appears to be a success story. Cloud computing has delivered extraordinary gains in scalability, efficiency and innovation. It has accelerated digital transformation and enabled entirely new categories of services.
However, beneath these gains lies an architectural paradox: The internet remains distributed in theory but is increasingly centralised in practice.
I believe that the most significant risk is not the cloud itself, but the hidden centralisation embedded within modern cloud architectures, particularly in the control layers that govern how systems discover, authenticate and interact with one another; when these layers fail, the impact is no longer localised, it becomes systemic.
Recent global outages have demonstrated that failures in seemingly small components such as DNS, identity systems or routing layers can cascade across regions, industries and services, disrupting everything from financial transactions to emergency services and traffic lights.
These events are not anomalies they are signals of a deeper structural issue. To understand this shift, we must move beyond traditional notions of infrastructure resilience and examine how modern systems fail, not just where they run, but how they are controlled.
In addition to exploring this shift, I introduce a practical framework for understanding cloud systemic risk and examines why the next generation of resilience strategies must account for a world where failure domains are no longer bounded by geography but by architecture.
Origins of the Internet
The internet is an independent public infrastructure without ownership or boundary. The infrastructure has become a critical backbone for digital services in today’s digital world, enabling business operations and public services worldwide; it underpins and straddles nearly all human activities in the modern age.
The name “Internet” derives from “internetworking” is an information superhighway made up of a mesh of disparate networks, it cannot be said to be owed by any entity, country or region.
The internet was decentralised both in terms of architecture and governance. The network was designed to withstand localised failures and ensure global connectivity, through globally dispersed routes, paths, domain name services etc ensuring its resilience and availability.
The internet’s evolution started as the commercialisation of the expanded ARPANET and the Academic network of networks in the 1990s. The commercialisation saw the entry of standalone internet service providers (ISPs) who provided connectivity and access to individuals, businesses and organisations, connecting them to the broader information superhighway via dial-up, DSL, cable, satellites, fibre etc.
Original Internet Architecture
In the early days, individual ISPs maintained autonomous systems (AS), configuring manual routes between known hosts, later evolving to dynamic routing with Internal Gateway Protocol (IGP) (for routing within the ISPs environment) and Border Gateway Protocol (BGP) (for routing to other ISPs). ISPs were classed and categorised on Tiers of 1 – 3.
Tier 1 ISPs: Global backbone providers that peer with each other.
Tier 2 ISPs: Regional providers that buy transit from Tier 1 and peer locally.
Tier 3 ISPs: Local ISPs serving end-users.
In the early 2000s several telecommunications services providers (telcos) became ISPs while some sold off the voice part of their business to fund their investment into Internet Protocol (IP) business.
These ISPs and telcos owned large data centres for hosting websites, enterprise apps and various connectivity services.
Internet exchange points (IXP) or Point of presence (PoP) are physical interconnection points where multiple ISPs and networks interconnect directly with each other to exchange traffic. They are located across major cities and towns across the world.
The Rise of Cloud
Contrary to conflicting definitions, the cloud is not the internet, while this confusion is understandable, it is important to call out the differences and similarities.
The concept and vision of “Cloud” was proposed in 1961 by MIT’s John McCarthy, who predicted computing would become a public utility. The word was used in the context of distributed computing as early as 1993 -1994 when Apple spinoff, General Magic and AT&T described “the Cloud” in discussions of programmable digital networks.
Key attributes of the cloud include;
- On-demand self-service: A consumer can unilaterally provision services such as computing (server time & network storage) and network access as needed automatically without requiring human interaction with a service provider.
- Broad network access: Capabilities are available over the network & accessed through standard mechanisms that promote use by heterogeneous thin or thick client platforms as well as other traditional or cloud-based software services.
- Rapid elasticity: Capabilities can be rapidly & elastically provisioned automatically to quickly scale out & rapidly released to quickly scale in. The available capabilities appear to be unlimited, in any quantity and at any time.
- Measured service: Cloud systems automatically control and optimise resource usage by using a metering capability at some level of abstraction appropriate to the type of service. Resource usage can be monitored, controlled & reported, providing transparency for both the provider & consumers.
- Resource pooling: Provider’s resources are pooled to serve multiple consumers using a multi-tenant model with different physical & virtual resources dynamically assigned & reassigned according to demand. A degree of location independence, users may be able to specify location e.g. Region or Country
Modern “Cloud computing” which started in early 2000s refers to the delivery of computing services such as servers, storage, databases, networking, software and analytics over the internet. It evolved partly as a natural evolution of internet utility and resource management. Instead of owning and maintaining physical hardware or data centres, users and organisations can access these resources on demand from cloud providers, paying only for what they use.
2002: Amazon launches early AWS services.
2006: AWS launches S3 and EC2, the true birth of modern cloud computing.
2008–2010: Google App Engine and Microsoft Azure launch, expanding the modern cloud ecosystem.
At the heart of cloud are data centres where physical computer resources, applications and other resources are pooled and virtualised which enables them to be packaged and sold as fractions of the overall pooled of resources. Typically, cloud services are categorised as compute, networking, storage, application services etc.
The integration of cloud services with the internet has now reached the point where one cannot easily be discerned from the other, creating a huge, converged network and infrastructure.
Modern “Cloud” Internet Architecture
Legacy internet was autonomous and decentralised without political or commercial ownership or control whereas the cloud is owned and managed by a handful of private business entities
What Has Changed?
Traditional Internet
- Thousands of ISPs
- Thousands of hosting providers
- Many DNS providers
- Distributed ownership
Modern Cloud Internet
- Few hyperscalers
- Shared platforms
- Shared identity systems
- Shared control planes
The internet itself has not become centralised, the services running on it have; cloud service providers may come and go but the internet will remain
Cloud Convergence & Evidence of Centralisation
At the start internet services, utilities and data centre hosting services were provided by thousands of commercial ISP organisations in different geographical regions of the world. These independent ISPs have since been acquired and merged by a handful of hyper scalers, hitherto referred to as cloud service providers globally. This convergence has significantly altered the decentralised and loosely governed nature of the internet as we knew it.
It is undeniable that this shift has delivered unprecedented scalability and convenience, but it has also introduced a critical vulnerability of “single points” of failure on a planetary scale due to convergence. Cloud service providers now own and control the IXPs/PoPs, routers, domain name servers and other internet utilities, although the access service are still being provided by telcos.
“According to estimates from Synergy Research Group, Amazon’s market share in the worldwide cloud infrastructure market amounted to 28 percent in the first quarter of 2026, ahead of Microsoft’s Azure platform at 21 percent and Google Cloud at 14 percent. Together, the “Big Three” hyperscalers account for more than 60 percent of the ever-growing cloud market, with the rest of the competition stuck in the low single digits”.
Worldwide market share of leading cloud infrastructure service providers in Q1 2026
Source: https://www.statista.com/
Cloud computing is no longer optional, it is foundational. Gartner forecasts global public cloud spending to exceed $700 billion in 2025, reflecting its central role in economic activity. [gartner.com]
At the enterprise level, cloud adoption has reached near saturation, with the majority of organisations running a large proportion of workloads in cloud environments. [datastackhub.com]
This concentration of workloads means that cloud platforms now underpin critical sectors, including:
- Financial systems
- Healthcare infrastructure
- Government services
- Communication platforms
- AI and data processing systems
As a result, cloud outages are no longer isolated technical events, they are systemic incidents.
This centralisation poses profound risk not just to businesses, but to critical national infrastructure and social stability. When a single provider falters, the ripple effect cascades across industries and global systems, disrupting essential services, financial systems and even critical national infrastructure from emergency services to financial services, hospitals, transportation networks, water, electricity and other utilities, reverberating through the fabric of modern society.
Whilst these platforms offer immense scalability and innovation, their central role introduces significant risks to societies, economies, governments and potentially human lives.
Despite improvements in infrastructure reliability, outages remain a persistent and costly risk. According to the Uptime Institute:
- Most organisations have experienced outages within the past three years
- More than two-thirds of outages cost over $100,000
- Complexity and third-party dependencies are increasing failure risks [datacenter…titute.com], [uptimeinstitute.com]
Modern cloud systems are highly interconnected, meaning failures rarely remain contained. Instead, they cascade across services, regions and industries.
The central thesis of this article is that modern cloud failures are increasingly architecture failures rather than infrastructure failures.
The Cloud Systemic Risk Model
Cloud concentration risk is often misunderstood as a simple infrastructure availability problem but in reality, systemic risk emerges from the interaction of four layers: the data plane, control plane, dependency layer, and concentration layer.
Layered cloud systemic risk model
In this model, I highlight that modern cloud failures increasingly originate in the control plane rather than the infrastructure itself. As dependencies accumulate and ownership becomes concentrated among a handful of hyperscale providers, failures within these layers can propagate well beyond their original point of impact.
Data Plane
Where applications, workloads and data operate, examples include Virtual machines, Databases, Containers, Storage systems and they typically fail locally.
Control Plane
DNS, identity & IAM systems, routing, API gateways, and orchestration services. A control-plane failure can affect systems that are otherwise healthy.
The greatest systemic risk resides in the control plane where failures can spread rapidly across otherwise resilient systems. Even if individual services are resilient, AWS, Azure, Google Cloud collectively host a huge proportion of the world’s digital workloads creating a concentration risk.
Dependency Layer
Modern services are no longer standalone, they consists of Microservices, APIs and SaaS integrations, one application may depend upon Azure, Microsoft 365, Salesforce, Stripe, Cloudflare, Okta etc. Failures propagate through these dependencies.
Concentration Layer
The ownership and control of infrastructure by a small number of hyperscalers. This is the strategic layer.
Outage Analysis (Case Study): 2020 – 2025
Major outages that demonstrated the fragility of cloud-dependent infrastructure
2025
AWS DNS Failure (October 2025)
This incident originated from a DNS resolution issue affecting DynamoDB service endpoints. Although initially localised to a single region, the failure propagated across multiple services due to dependencies on shared control-plane functions.
- Root cause: DNS/control-plane failure
- Impact: EC2, Lambda, S3, and multiple dependent applications
- Outcome: Global service disruption across finance, media, and enterprise systems
The incident demonstrated that control-plane dependencies can override regional isolation assumptions. [axisops.com]. What makes this incident particularly significant is that the underlying infrastructure remained largely operational. The outage stemmed from a control-plane dependency, reinforcing the argument that modern cloud failures are increasingly governance and orchestration failures rather than infrastructure failures.
Azure Front Door Outage (29 October 2025)
Nine days after the AWS incident, Azure experienced a global outage linked to a misconfiguration in its Front Door service.
- Root cause: Routing and configuration error
- Impact: Microsoft 365, enterprise SaaS platforms
- Outcome: Global service degradation
This case highlights how edge routing and global entry points represent critical failure domains. [windowsforum.com]
Google Cloud Networking Incident (June 2025)
A configuration change triggered a multi-region service disruption affecting compute and AI workloads.
- Root cause: Network configuration error
- Impact: Cross-regional services and APIs
- Outcome: Enterprise application downtime
Cloudflare DNS Outage (July 2025)
Cloudflare’s public DNS resolver experienced disruption due to a topology change.
- Root cause: Network policy/configuration change
- Impact: Internet-wide name resolution degradation
- Outcome: Widespread accessibility issues
2024
CrowdStrike / Microsoft Ecosystem Failure (July 2024)
Although not purely a cloud provider outage, a faulty update affected millions of Windows systems, demonstrating systemic dependency on shared digital infrastructure.
- Root cause: Software update error
- Impact: Airlines, hospitals, financial systems
- Estimated economic impact: >$5 billion
This event illustrates how tightly coupled ecosystems amplify failure effects. [linkedin.com]
Google Cloud:
Frankfurt power failure caused hours of downtime for European customers.
AWS
Multiple outages totalling ~100 hours downtime globally; six outages lasted over 10 hours (source: cyberinsur…cenews.org)
2023
- AWS US-East-1: June – Extended outage estimated to cost $3.4B for 24 hours if prolonged; $7.8B for 48 hours. (source: crn.com)
- Azure: March – Identity service failure disrupted authentication for global users.
2022
- Google Cloud: November – Networking issue impacted multiple regions, affecting enterprise workloads.
- AWS: June – Power outage in a data centre caused EC2 and RDS failures in US-East-2.
2021
- AWS Global: December – Network congestion caused widespread downtime for Netflix, Disney+ and Amazon services.
- Azure: April – DNS outage disrupted Microsoft 365 and Teams worldwide.
2020
- AWS US-East-1: November – Kinesis service failure disrupted major apps like Roku and Adobe for hours.
- Google Cloud: December – Authentication outage affected Gmail, YouTube and Google Workspace globally.
Implication for Critical Sectors
Healthcare
Cloud outages can delay access to patient records, disrupt scheduling systems, and impact emergency care.
Finance & Economic Disruption
Banking platforms, payment systems, and trading environments rely on real-time data availability, making them highly sensitive to downtime. Downtime can freeze markets and halt commerce. Businesses lose revenue, supply chains stall and consumers face service interruptions.
Government and Social Order
Cloud-hosted identity systems and public services introduce national security and governance risks during outages. Governance and societal stability are increasingly intertwined with digital infrastructure. Modern military communications, intelligence analysis and logistics frequently leverage commercial cloud platforms. The concentration of these services raises several concerns:
- Strategic Vulnerability: An outage could impede command and control systems, intelligence sharing, or the operation of autonomous defence technologies.
- National Security Risks: Foreign adversaries may target cloud infrastructure to disrupt critical national capabilities.
- Social Order: Public services and emergency response systems may falter during outages, eroding trust in institutions and exacerbating social unrest during crises.
- Messaging apps, collaboration tools and social platforms go dark, cutting off personal and professional communication.
AI Systems
AI workloads are increasingly centralised in cloud-hosted GPU clusters. Disruptions can halt inference, degrade models, and impact autonomous systems. Autonomous AI systems ranging from self-driving vehicles to automated trading engines and smart city management are heavily reliant on cloud resources for computation, data storage and coordination. Outages or performance degradation in cloud services can cause:
- Operational Paralysis: Without access to critical data or processing power, AI-driven systems may halt unexpectedly, leading to safety incidents or financial losses.
- Data Integrity Risks: Interrupted synchronisation or corrupted data during outages could degrade the performance and trustworthiness of AI models.
- Security Vulnerabilities: Failover to less-secure backup systems during outages may create exploitable gaps for malicious actors.
The dependency of autonomous systems on a handful of cloud providers thereby amplifies the scale and potential impact of any disruption.
Critical National Infrastructure
- Energy Grids: Cloud-based control systems for electricity distribution and smart grids depend on real-time data processing. A major outage could interrupt monitoring, coordination and even blackout response.
- Transport and Logistics: Airlines, railways and supply chains rely on cloud-managed scheduling and tracking, risking widespread delays and economic losses if access is interrupted.
These examples highlight the systemic nature of concentration risk: a single point of failure can have far-reaching, real-world consequences
Root Causes of Modern Cloud Failures
In recent years, high-profile outages involving Azure, AWS and GCP have underscored the fragility of this concentrated model, importantly, many failures originate not from hardware faults, but from software and operational processes in highly complex distributed systems.
These incidents have often been triggered by:
- Configuration errors and human mistakes
Human error & mistakes in system administration or incident response can exacerbate the duration and scope of outages accounting for 66–68% of outages in 2024. (cyberinsur…cenews.org), (cxotoday.com). This is a leading cause of outages [productivehub.com].
- Software bugs and faulty updates
Software Updates Gone Wrong: Faulty patches or configuration errors can propagate rapidly across global data centres, causing widespread outages. Change management errors account for 42% of incidents (rollouts, config updates). (datastackhub.com)
- Network and control-plane failures
Disruptions in backbone connectivity or routing errors can isolate entire regions from cloud services. Networking/control-plane issues: 31%. (datastackhub.com)
- Infrastructure complexity and interdependencies
Unexpected surges in demand or denial-of-service attacks can overwhelm capacity, leading to service degradation or downtime.
Emerging Systemic Risks
The convergence of cloud services creates a new category of risk:
- Single-provider dependency
- Hidden control-plane centralisation
- Global failure domains
- Interconnected service dependencies
As Uptime Institute highlights, modern outages increasingly result from complex system interactions rather than isolated failures. [datacentre.solutions]
This complexity makes failures harder to predict, contain, and recover from.
Future Risk Outlook
Several trends are likely to increase systemic cloud risk over the next decade, in my view these include:
AI Centralisation
Training and inference workloads are increasingly concentrated in hyperscale GPU clusters.
Digital Infrastructure Dependency
Critical sectors including healthcare, energy and transportation continue to migrate to cloud-based platforms.
Increasing Interconnectivity
Modern applications depend on hundreds of interconnected services that amplify outage propagation.
Sovereignty and Geopolitical Risk
National dependence on foreign-owned cloud infrastructure raises strategic and regulatory challenges.
The combination of these trends means future outages may have greater societal and economic consequences than those experienced today.
Mitigation Strategies
My recommendations to address these risks, is that organisations must adopt resilience-first architectures through
- Multi-Cloud Strategies
Distribute workloads across providers to reduce dependency.
- Active-Active Architectures
Avoid cold failover; ensure systems run across regions/providers simultaneously.
- Decoupled Identity & DNS
Reduce reliance on single-provider control-plane services.
- Chaos Engineering & Resilience Testing
Regularly test failure scenarios to validate resilience.
- Regulatory Oversight
Treat cloud platforms as critical infrastructure requiring resilience standards.
- Edge computing to reduce central dependence
What Most Organisations Get Wrong About Cloud Resilience
Many organisations believe that deploying workloads across multiple availability zones automatically delivers resilience. In practice, resilience is often constrained by shared control planes, identity systems, DNS services and operational processes. The lesson is that redundancy does not necessarily eliminate concentration risk.
Limitations
My analysis relies primarily on publicly reported outage events and industry research; the true frequency, causes and impacts of cloud failures may be underreported due to limited disclosure by providers. Furthermore, cloud infrastructure remains highly reliable relative to many traditional enterprise environments. The focus of this article is therefore on systemic concentration risk rather than overall reliability.
Conclusion
Cloud computing has transformed the internet, enabling innovation at unprecedented scale. However, this transformation has come with a trade-off: increased systemic risk due to centralisation.
I believe that technology leaders, policymakers and risk managers must collaborate to diversify digital supply chains, enforce resilience standards and ensure that the backbone of our digital future is robust, adaptable and secure.
The future of digital resilience depends on recognising that efficiency and resilience must be balanced. Without deliberate architectural, organisational, and regulatory interventions, the same forces that enabled the cloud revolution may also become its greatest vulnerability



