Summary
- A utility fault in the Netherlands was followed by cooling disruption inside Google Cloud’s europe-west4-a zone.
- Rising temperatures led Google to shut down selected workloads to protect physical equipment and customer data.
- Recovery depended on both electrical and mechanical systems, while customers without pre-built regional redundancy had limited options.
A utility power failure at a Dutch facility serving Google Cloud developed into a cooling outage that forced engineers to shut down selected infrastructure services in the europe-west4-a zone.
Google Cloud VMware Engine, Bare Metal Solution, and Google Cloud NetApp Volumes were affected. Temperature alarms began during the evening of 15 July, after which workloads were progressively turned down as conditions inside the facility deteriorated.
Google said the shutdowns protected customer data and physical hardware from excessive temperatures. VMware customers lost connectivity, managed NetApp clusters were taken offline, and bare-metal servers were powered down, with the company advising customers to route traffic elsewhere where multi-region capacity was already available. Its incident record shows more than 12 hours of material disruption, followed by residual work on a small number of bare-metal systems.
Electrical autonomy ran into a thermal limit
Once the utility disturbance affected cooling, the safe operating period was governed by temperature rather than the amount of backup electricity available. Servers, storage, networking, UPS equipment, and power-conversion plant continued producing heat, and the facility could not sustain those loads indefinitely without working pumps, chillers, valves, controls, or heat rejection.
The available thermal ride-through depends on rack density, airflow, room volume, chilled-water inventory, pipework, control behaviour, and the distribution of load across the hall. High-density equipment reduces the time available to diagnose a cooling fault, particularly where several racks concentrate a large amount of power into a small area.
Google identified an upstream utility fault as the initiating event, followed by failures involving electrical distribution and cooling. Its public record does not show that every critical system failed simultaneously. Power incidents often propagate through dependencies, affecting control panels, pumps, valves, cooling towers, or heat exchangers even where generators and UPS systems continue carrying part of the site.
Protective shutdowns prevented further thermal exposure, although they created a long restoration sequence. Storage and virtualised environments had to return in a controlled order, network paths needed to be re-established, and customer systems required checks before normal operation resumed.
Bare-metal services create particular constraints because workloads are tied to physical machines in a specific zone. They cannot always be recreated elsewhere as quickly as a virtual resource, especially where the alternative facility lacks spare hardware or an identical configuration.
Cloud resilience stops at the facility boundary
Availability zones can reduce exposure to individual buildings or systems, but the separation varies by service. A customer using a zonal managed service remains dependent on the power, cooling, network, and operational conditions of that location, while a multi-region architecture requires data replication, network capacity, compatible services, and recovery procedures long before an incident begins.
Google’s advice to redirect traffic could only help customers that had already provisioned another location. Identity controls, routing, replicated data, spare compute, and tested automation all carry additional cost, which leaves some deployments concentrated in a single zone despite the availability of wider options.
The incident also joins cloud continuity to physical plant. Software orchestration can move an application only when another powered and cooled environment is ready to accept it. Every layer of abstraction still depends on switchgear, generators, pumps, cooling equipment, controls, cabling, and people.
Engineers examining the facility will need to establish how the utility event propagated, which cooling components became unavailable, how the control system responded, and how long the halls remained inside their safe thermal envelope. Google has not published an engineering root-cause report with that level of detail.
Power restoration alone did not complete the recovery. Cooling had to return before workloads could be restarted safely, after which storage, virtualisation, networking, and physical servers moved through separate restoration processes.
That sequence places mechanical systems alongside generators and UPS equipment in the resilience hierarchy. A site with substantial electrical autonomy can still lose usable capacity when heat cannot be removed, while a cooling design with strong redundancy may remain exposed if its controls or pumps share a vulnerable power path.
The next useful disclosure would describe the affected components, the transfer sequence, the thermal ride-through period, and the corrective work. Customers currently have a detailed service timeline, but not enough facility information to judge whether the same dependency has been removed.

