{
  "id": 13145801,
  "title": "Growing pains: how distributed AI training changes the network between datacenters",
  "url": "https://urgent.news/2026/10/09/growing-pains-how-distributed-ai-training-changes-the-network-between-13145801",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-09T15:00:00.000Z",
  "source": {
    "name": "The Register Science",
    "slug": "the-register-science",
    "url": "https://www.theregister.com/networks/2026/10/09/sponsored-growing-pains-how-distributed-ai-training-changes-the-network-between-datacenters/5301554"
  },
  "original_language": "en",
  "account": "As AI models grow in complexity and size, the need for distributed training across multiple datacenters has become essential. This shift is driven by the limitations of single datacenters or campus clusters, which can struggle to meet the high compute and power demands of new models. The largest frontier training runs could consume anywhere from 4-16GW of power by 2030, necessitating distributed training to find power and space for construction and avoid planning constraints.\n\nTraining a large language model (LLM) involves a neural network with billions of numerical parameters, or weights, that are adjusted during training. Thousands of GPUs and other accelerators work together to process batches of data and update weights. However, they cannot operate independently, as they must regularly synchronize and exchange results before progressing.\n\nThe distribution of compute across multiple sites poses challenges, particularly in terms of synchronization. Traditional datacenter interconnects (DCI) are designed for redundancy, reach, and workload distribution, but their asynchronous nature can lead to issues when used for distributed AI training. These issues include bursty traffic, incast problems, and the potential for the network to become a bottleneck, causing additional latency and delaying the training process.\n\nTo address these challenges, Cisco's Itamar Gold suggests that short-range datacenter optics are not suited for long-range, high-bandwidth connections required for scale-across AI training. Instead, traffic should be funneled through narrower pipes to avoid becoming a choke point during synchronized bursts. While the network may experience occasional congestion or packet loss, these issues should not halt the entire training process. However, serious failures may require training to restart from a saved checkpoint.\n\nIn summary, the growth of distributed AI training is a response to the increasing demands of large language models and the limitations of single datacenters. While distributed training brings new challenges in terms of synchronization, bandwidth requirements, and network architecture, careful planning and the use of specialized interconnects can help mitigate these issues and enable efficient and effective AI model training across multiple datacenters.",
  "summary": "SPONSORED FEATURE: While linking the GPUs in a datacenter is a challenge, coordinating thousands of them to work efficiently over hundreds of kilometers presents a new set of difficulties.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "The Register",
        "title": "Growing pains: how distributed AI training changes the network between datacenters",
        "url": "https://urgent.news/2026/10/09/growing-pains-how-distributed-ai-training-changes-the-network-between",
        "published": "2026-10-09T15:00:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}