When one datacenter is no longer enough
SPONSORED POST: Why training AI across multiple sites is creating a new networking challenge
Training massive AI models has evolved far beyond simply adding more GPUs. As clusters expand in size, the availability of power becomes a daunting physical limitation. This necessitates infrastructure teams to contemplate how a single training workload can function across multiple data centers. This scenario presents a vastly different networking challenge.
Traditional datacenter interconnect was intended for moving traffic between sites. However, AI training requires a much more precise approach: enormous, synchronized data flows with minimal packet loss and tightly coordinated communication between GPUs. If one segment of the cluster encounters a slowdown, its effects can propagate to the entire job.
In this interview, The Register's Tim Phillips engages in conversation with Rakesh Chopra, Cisco's Senior Vice President for Silicon and Systems Architecture and a Cisco Fellow, to delve into the "Scale-Across" imperative and discuss why Cisco advocates for a novel approach to routing in distributed AI infrastructure.
Rakesh explains how the demands of AI are transforming the role of the network, shifting it from mere transport to an integral component of the compute system. He discusses the challenge of making geographically dispersed facilities operate as a single, coherent machine. He emphasizes the importance of buffering, high-speed coherent optics, and tightly integrated silicon when workloads traverse lengthy fiber links.
The conversation also delves into the intricacies of the cluster itself. Rakesh outlines how Cisco's Silicon One architecture and Intelligent Collective Networking are engineered to handle synchronized surges of GPU traffic, mitigate bottlenecks, and expedite job completion time. Power efficiency is a crucial aspect of this discussion, encompassing the trade-off between network consumption and the energy allocated to GPUs.
Security and longevity are also addressed. The interview explores hardware-accelerated MACsec and IPsec, as well as the role of programmability in enabling infrastructure to adapt to the ever-evolving landscape of AI workloads. For infrastructure and datacenter leaders planning for more expansive, distributed AI environments, this discussion offers a pragmatic perspective on the networking hurdles that arise when a single site is no longer sufficient.
Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- When one datacenter is no longer enough theregister.com