Urgent.News

What's breaking now, across thousands of outlets.

AI

Growing pains: how distributed AI training changes the network between datacenters

SPONSORED FEATURE: While linking the GPUs in a datacenter is a challenge, coordinating thousands of them to work efficiently over hundreds of kilometers presents a new set of difficulties.

Growing pains: how distributed AI training changes the network between datacenters

As AI models grow in complexity and size, the need for distributed training across multiple datacenters has become essential. This shift is driven by the limitations of single datacenters or campus clusters, which can struggle to meet the high compute and power demands of new models. The largest frontier training runs could consume anywhere from 4-16GW of power by 2030, necessitating distributed training to find power and space for construction and avoid planning constraints.

Training a large language model (LLM) involves a neural network with billions of numerical parameters, or weights, that are adjusted during training. Thousands of GPUs and other accelerators work together to process batches of data and update weights. However, they cannot operate independently, as they must regularly synchronize and exchange results before progressing.

The distribution of compute across multiple sites poses challenges, particularly in terms of synchronization. Traditional datacenter interconnects (DCI) are designed for redundancy, reach, and workload distribution, but their asynchronous nature can lead to issues when used for distributed AI training. These issues include bursty traffic, incast problems, and the potential for the network to become a bottleneck, causing additional latency and delaying the training process.

To address these challenges, Cisco's Itamar Gold suggests that short-range datacenter optics are not suited for long-range, high-bandwidth connections required for scale-across AI training. Instead, traffic should be funneled through narrower pipes to avoid becoming a choke point during synchronized bursts. While the network may experience occasional congestion or packet loss, these issues should not halt the entire training process. However, serious failures may require training to restart from a saved checkpoint.

In summary, the growth of distributed AI training is a response to the increasing demands of large language models and the limitations of single datacenters. While distributed training brings new challenges in terms of synchronization, bandwidth requirements, and network architecture, careful planning and the use of specialized interconnects can help mitigate these issues and enable efficient and effective AI model training across multiple datacenters.

Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in AI

Author or artificial intelligence? Literature in crisis

AI is a hot topic at this year's Frankfurt Book Fair. An international literary scandal raises questions about machine-generated texts.

  • AI's impact on literature sparks debates at Frankfurt Book Fair
  • Thélyson Orélien's novel accused of AI authorship, sparking controversy
  • Authors and publishers call for stricter AI regulation and copyright protection

More from Friday 9 October →