Nebius resolves GPU interconnect failure

Nebius resolves GPU interconnect failure

Nebius suffered InfiniBand connectivity failures across GPU clusters on Saturday.

Nebius resolves GPU interconnect failure
Summary
  • Some Nebius GPU clusters in eu-west2 suffered InfiniBand connectivity problems from 02:30 to 08:59 UTC on 3 October.
  • Distributed training jobs could experience degraded interconnect performance or failures during the incident.
  • Nebius identified the cause at 04:45 UTC, deployed a fix at 07:44, and declared the incident resolved at 08:59.

Nebius has resolved an InfiniBand connectivity incident that affected some GPU clusters in its eu-west2 region and created a risk of degraded or failed distributed AI training workloads.

The AI cloud provider began investigating the problem at 02:30 UTC on Saturday, 3 October. Its status page said affected distributed training jobs could experience degraded interconnect performance or fail while the issue remained active.

Nebius said it had identified the cause by 04:45 UTC, with an on site team working on the fault. A fix was implemented at 07:44 and moved into monitoring before the company declared the incident resolved at 08:59.

The status update did not disclose the underlying technical cause, the number of clusters affected, or whether any customer training jobs had to be restarted. It also did not report a general loss of the eu-west2 region.

The failure is significant because InfiniBand is not an ancillary network in large AI clusters. GPU training depends on accelerators exchanging large volumes of data with very low latency. When that fabric degrades, the compute nodes themselves may remain powered and available while the workload running across them becomes inefficient or unusable.

Nebius markets its AI cloud around large GPU clusters using NVIDIA accelerators and InfiniBand networking. That architecture allows customers to scale training across many GPUs, but it also means network performance forms part of the compute platform rather than sitting outside it.

A conventional server application may tolerate a temporary reduction in network throughput without failing outright. Distributed model training is less forgiving because accelerators repeatedly synchronise work across nodes. A slow or unavailable interconnect can hold back the entire job, leaving costly GPU capacity waiting for communications to recover.

The incident therefore highlights a different resilience metric from basic server uptime. For AI infrastructure, operators increasingly have to measure whether the cluster can sustain the latency, bandwidth, and collective communications performance expected by the workload, not merely whether individual hosts respond.

Nebius has published performance and reliability metrics for its GPU infrastructure and promotes resilient operation as a feature of the platform. Saturday’s event does not by itself establish a broader reliability problem, but it shows why the interconnect fabric has to be treated as critical infrastructure within a GPU cluster.

The company’s recent status history also includes separate service incidents, including network problems in eu-west2 in September and a short period of partial virtual machine creation and object storage unavailability on 2 October. Nebius has not linked those events to Saturday’s InfiniBand fault.

With the latest incident closed, the remaining question is whether Nebius publishes further technical detail about the failure. For customers running long model training jobs, information about fault domains, recovery behaviour, and whether jobs need manual intervention can be as important as the headline duration of an outage.


Stay updated with the latest insights and trends in the data centre industry by subscribing to our newsletter.

← Back

Thank you for your response. ✨