Stakefish outage exposes London power-path weakness

Stakefish outage exposes London power-path weakness

Stakefish says a planned maintenance event at a London data centre triggered a power loss affecting more than 6,000 blockchain validators for over 21 hours.

Stakefish outage exposes London power-path weakness
Summary
  • Planned electrical work left part of the infrastructure operating on a single distribution path before a circuit breaker tripped.
  • More than 6,000 Lido validators were affected, with full recovery taking around 21 hours and 36 minutes.
  • Stakefish identified weak failover design, equipment concentration, and inadequate real-time communication with the facility as contributing factors.

Planned electrical maintenance at a London data centre triggered a power failure that affected more than 6,000 Lido validators operated by Stakefish, according to a post-incident report published by the infrastructure provider.

The August incident lasted around 21 hours and 36 minutes from the initial service impact to full recovery and exposed weaknesses in both the facility power path and Stakefish’s own high-availability design.

During planned upstream maintenance, the A power distribution line serving the relevant infrastructure was isolated for electrical engineering work. Stakefish’s report refers inconsistently to the affected location as a Latitude data centre and as Telehouse South (LON2), so DataCentral is retaining those source labels rather than resolving the facility attribution independently.

That left the power distribution unit supplying high-density validator cabinets operating from a single line. A circuit breaker subsequently tripped, causing multiple physical servers to lose power.

The outage began at 00:04 UTC on 8 August. Stakefish said service recovery started at 05:40 UTC and more than 5,000 validators were back by 10:04 UTC, but 1,286 remained offline until full recovery at 21:40 UTC.

The company later compensated users with 6.9613 ETH for losses linked to the incident.

The electrical failure was only one layer of the event. Stakefish’s post-mortem identifies four broader problems: a lack of direct real-time feedback from the data centre, a decision to wait for facility recovery instead of failing over, insufficient high-availability design within the validator infrastructure, and concentration of physical servers in a single cabinet.

Those findings make the incident a useful example of why nominal data centre redundancy does not automatically create application resilience.

Maintenance is one of the periods when redundant infrastructure is intentionally reduced. A system designed around two independent paths may temporarily operate on one while switchgear, distribution, UPS equipment, or other components are isolated. During that window, an otherwise tolerable second failure can become service-affecting.

Operators therefore need maintenance procedures that account for the changed fault domain, while customers need to understand what facility work does to the assumptions behind their own architecture.

Stakefish’s account shows that its application-level response was also constrained by incomplete information. The company said its team did not receive timely updates on the progress of facility recovery and made decisions based partly on previous experience rather than current conditions.

That delayed the decision to move workloads elsewhere. In infrastructure designed for continuous service, the communication path between a data centre and a customer can therefore become part of the resilience system rather than an administrative extra.

The concentration of servers in one cabinet added another single point of failure. Even where the building offers resilient electrical infrastructure, placing a large proportion of a service into one distribution domain can defeat much of that redundancy.

Stakefish said its validator architecture uses dedicated bare-metal servers and remote signing infrastructure, with monitoring through Prometheus, Grafana, and Alertmanager. The incident showed that monitoring can identify a failure without necessarily providing the physical-site information required to choose the right recovery action.

The post-mortem points towards a familiar operational lesson: redundancy has to extend from the utility connection through facility distribution and customer equipment to application failover. If any one of those layers remains concentrated, a maintenance event can expose it.

Stakefish’s published corrective actions will now be measured against whether future maintenance or single-cabinet failures can be contained without repeating a multi-hour loss of service.


Stay updated with the latest insights and trends in the data centre industry by subscribing to our newsletter.

← Back

Thank you for your response. ✨