Proton outage exposes cooling failover weakness

Proton outage exposes cooling failover weakness

A Frankfurt cooling failure knocked multiple Proton services temporarily offline.

Proton outage exposes cooling failover weakness
Summary
  • Proton suffered a global service outage after a critical cooling failure at its Frankfurt data centre.
  • Temperatures reportedly rose from roughly 30°C to 60°C in about 20 minutes during the failure.
  • Automatic failover did not behave as expected because the Frankfurt environment was only partially unavailable, forcing manual recovery.

Proton has restored its services after a critical cooling failure at a Frankfurt data centre triggered a global outage and exposed a weakness in the company’s automatic failover process.

Proton began investigating the incident shortly after midnight CEST on 27 August. Its status page identified a critical cooling failure in Frankfurt at 00:38 and said traffic was being shifted to backup sites. Most services were recovering by 01:57, with the remaining minor issues reported resolved later in the incident.

The company said no data was lost, although some users experienced delays in email delivery, reception, and notifications during recovery. Proton has said it will publish a post-mortem after completing its incident clean-up.

The physical behaviour of the facility was unusually visible because Proton founder and chief executive Andy Yen subsequently published temperature information from the event. According to Yen, the environment moved from a baseline of around 30°C to a peak of approximately 60°C in roughly 20 minutes after the cooling failure.

That rate of temperature increase demonstrates how little thermal margin a heavily loaded data hall can have once active heat removal stops. Servers continue converting electrical power into heat until they shut down, throttle, or lose power, while the thermal mass of the room and cooling system only delays the rise.

Proton’s recovery was complicated by the fact that the Frankfurt environment did not fail completely. Yen said the site was only partially unavailable, which meant the company’s automatic failover did not activate as expected. Engineers initially attempted to restore cooling before the servers suffered permanent damage, slowing the transition to backup locations.

Partial failures are harder than clean failures

That sequence is an operational-resilience problem rather than simply a cooling-equipment problem. Disaster-recovery systems are often easiest to automate when a failure produces an unambiguous condition: a site disappears, a network path is lost, or an entire service becomes unavailable.

A degraded site can be more difficult. Some servers, links, or control systems may continue to respond while environmental conditions are deteriorating. Automation then has to decide whether to leave workloads in place, evacuate them, shut systems down, or wait for human intervention.

Proton said the particular failure mode was known but that mitigations had not been prioritised because the scenario was judged extremely unlikely. The incident will now force that risk assessment to be revisited.

The company has not identified the Frankfurt data centre operator or disclosed the precise cooling fault. There is therefore no basis yet to determine whether the event began with a chiller, pump, control, power, heat-rejection, or other mechanical failure. The promised post-mortem will be important in separating the initial physical fault from the subsequent resilience and traffic-management issues.

Frankfurt is strategically useful to Proton because of its network connectivity. The company has said the location was added partly for bandwidth and proximity to one of Europe’s largest internet-exchange ecosystems, while its Swiss facilities in Geneva and Zurich remain in service.

That geographic diversity provided alternative capacity during the outage, but the incident shows that having backup sites is only one layer of resilience. Workloads also need sufficiently fast detection, decision logic, replication, network routing, and operating procedures to move away from a failing environment before the physical condition damages equipment or stretches service-level targets.

Cooling failures are particularly unforgiving because the escalation time can be measured in minutes rather than hours. Higher rack densities increase the amount of heat concentrated into the same room volume, making detection, control integration, and orderly workload evacuation more important as facilities host more power-intensive compute.

Proton’s systems were operational again by 28 August, according to its status page. The next test will be whether its post-mortem identifies both the root cooling fault and the reason a known partial-failure scenario remained insufficiently mitigated before the incident occurred.


Stay updated with the latest insights and trends in the data centre industry by subscribing to our newsletter.

← Back

Thank you for your response. ✨