{
  "id": 13140967,
  "title": "Growing pains: how distributed AI training changes the network between datacenters",
  "url": "https://urgent.news/2026/10/09/growing-pains-how-distributed-ai-training-changes-the-network-between",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-09T15:00:00.000Z",
  "source": {
    "name": "The Register",
    "slug": "the-register",
    "url": "https://www.theregister.com/networks/2026/10/09/sponsored-growing-pains-how-distributed-ai-training-changes-the-network-between-datacenters/5301554"
  },
  "original_language": "en",
  "account": "Distributed AI training is expanding beyond individual data centers, according to recent developments. Major cloud providers and tech giants like Google, Microsoft, AWS, Meta, CoreWeave, and Google Cloud have connected AI compute clusters across various locations to facilitate the training of large language models (LLMs). This trend is driven by the increasing computational power and power demands of new models, which often exceed the capacity of single data centers. Cisco estimates that training some models may require clusters with tens of thousands of GPUs, while a few 2030 frontier training runs could draw 4-16GW of power. Distributing compute resources in this manner allows for greater flexibility in selecting locations with ample power, space, and fewer planning constraints.\n\nTraining an LLM involves a neural network with billions of weights that are adjusted during the training process. The training data is divided into batches, which are processed by thousands of GPUs in parallel. However, this distributed approach necessitates synchronization between the GPUs to exchange results and update weights before proceeding to the next stage. This synchronization process generates bursty traffic throughout the training, potentially causing congestion and \"incast\" problems where multiple senders attempt to communicate to the same destination simultaneously.\n\nTo optimize this distributed training process, Cisco's Ramesh Sivakolundu suggests enhancing network performance by reducing synchronization delays and avoiding bottlenecks. In a more constrained inter-site fabric used for scale-across AI training, traffic is funneled through a narrower pipe, potentially leading to choke points. Traditional data center interconnects (DCI) are designed for redundancy and workload distribution, while the AI fabric requires sufficient bandwidth and predictable delivery to prevent delays or dropped flows from causing additional latency.",
  "summary": "SPONSORED FEATURE: While linking the GPUs in a datacenter is a challenge, coordinating thousands of them to work efficiently over hundreds of kilometers presents a new set of difficulties.",
  "key_points": [
    "Distributed AI training expands beyond single data centers.",
    "Major cloud providers and tech giants connect AI clusters across locations.",
    "Cisco recommends optimizing network performance for distributed training."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "The Register Science",
        "title": "Growing pains: how distributed AI training changes the network between datacenters",
        "url": "https://urgent.news/2026/10/09/growing-pains-how-distributed-ai-training-changes-the-network-between-13145801",
        "published": "2026-10-09T15:00:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}