When one datacenter is no longer enough
SPONSORED POST: Why training AI across multiple sites is creating a new networking challenge
Training the most powerful AI models has evolved beyond simply adding more GPUs. As these clusters expand, power availability becomes a physical constraint, prompting infrastructure teams to explore operating a single training workload across multiple data centers. This shift introduces a new networking challenge that traditional datacenter interconnects were not designed to handle.
AI training demands precise, synchronous flows with minimal packet loss and tightly coordinated communication between GPUs. If one segment of the cluster slows down, the impact can spread throughout the entire job.
In this interview with The Register's Tim Phillips, Rakesh Chopra, SVP, Silicon and Systems Architecture, Cisco Fellow, discusses the "Scale-Across" imperative that Cisco promotes. Cisco believes that distributed AI infrastructure requires a new approach to routing, transforming the network from a simple transport mechanism to an integral part of the compute system.
Rakesh explains the challenges of making geographically separated facilities work together as a single, deterministic machine. He emphasizes the importance of buffering, high-speed coherent optics, and tightly integrated silicon when workloads stretch across long-distance fiber links. The conversation also delves into the cluster's inner workings, highlighting how Cisco's Silicon One architecture and Intelligent Collective Networking are designed to manage synchronized GPU traffic, mitigate bottlenecks, and shorten job completion times.
Power efficiency is a critical aspect, with Cisco considering the trade-off between network consumption and the energy available to GPUs. The interview also touches upon security and longevity, including hardware-accelerated MACsec and IPsec, as well as the role of programmability in helping infrastructure adapt to the evolving AI workloads.
For infrastructure and datacenter leaders planning for larger, more distributed AI environments, this discussion provides a practical view of the networking challenges that arise when one site is no longer sufficient.
Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
Also reported by 1 other outlet
- When one datacenter is no longer enough theregister.com