Urgent.News

What's breaking now, across thousands of outlets.

AI

Growing pains: how distributed AI training changes the network between datacenters

SPONSORED FEATURE: While linking the GPUs in a datacenter is a challenge, coordinating thousands of them to work efficiently over hundreds of kilometers presents a new set of difficulties.

Growing pains: how distributed AI training changes the network between datacenters

Distributed AI training is expanding beyond individual data centers, according to recent developments. Major cloud providers and tech giants like Google, Microsoft, AWS, Meta, CoreWeave, and Google Cloud have connected AI compute clusters across various locations to facilitate the training of large language models (LLMs). This trend is driven by the increasing computational power and power demands of new models, which often exceed the capacity of single data centers.

Cisco estimates that training some models may require clusters with tens of thousands of GPUs, while a few 2030 frontier training runs could draw 4-16GW of power. Distributing compute resources in this manner allows for greater flexibility in selecting locations with ample power, space, and fewer planning constraints.

Training an LLM involves a neural network with billions of weights that are adjusted during the training process. The training data is divided into batches, which are processed by thousands of GPUs in parallel. However, this distributed approach necessitates synchronization between the GPUs to exchange results and update weights before proceeding to the next stage.

This synchronization process generates bursty traffic throughout the training, potentially causing congestion and "incast" problems where multiple senders attempt to communicate to the same destination simultaneously.

To optimize this distributed training process, Cisco's Ramesh Sivakolundu suggests enhancing network performance by reducing synchronization delays and avoiding bottlenecks. In a more constrained inter-site fabric used for scale-across AI training, traffic is funneled through a narrower pipe, potentially leading to choke points.

Traditional data center interconnects (DCI) are designed for redundancy and workload distribution, while the AI fabric requires sufficient bandwidth and predictable delivery to prevent delays or dropped flows from causing additional latency.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at theregister.com →

More in AI

Author or artificial intelligence? Literature in crisis

AI is a hot topic at this year's Frankfurt Book Fair. An international literary scandal raises questions about machine-generated texts.

  • AI's impact on literature sparks debates at Frankfurt Book Fair
  • Thélyson Orélien's novel accused of AI authorship, sparking controversy
  • Authors and publishers call for stricter AI regulation and copyright protection

More from Friday 9 October →