Urgent.News

What's breaking now, across thousands of outlets.

AI

When one datacenter is no longer enough

SPONSORED POST: Why training AI across multiple sites is creating a new networking challenge

When one datacenter is no longer enough

Training massive AI models has evolved far beyond simply adding more GPUs. As clusters expand in size, the availability of power becomes a daunting physical limitation. This necessitates infrastructure teams to contemplate how a single training workload can function across multiple data centers. This scenario presents a vastly different networking challenge.

Traditional datacenter interconnect was intended for moving traffic between sites. However, AI training requires a much more precise approach: enormous, synchronized data flows with minimal packet loss and tightly coordinated communication between GPUs. If one segment of the cluster encounters a slowdown, its effects can propagate to the entire job.

In this interview, The Register's Tim Phillips engages in conversation with Rakesh Chopra, Cisco's Senior Vice President for Silicon and Systems Architecture and a Cisco Fellow, to delve into the "Scale-Across" imperative and discuss why Cisco advocates for a novel approach to routing in distributed AI infrastructure.

Rakesh explains how the demands of AI are transforming the role of the network, shifting it from mere transport to an integral component of the compute system. He discusses the challenge of making geographically dispersed facilities operate as a single, coherent machine. He emphasizes the importance of buffering, high-speed coherent optics, and tightly integrated silicon when workloads traverse lengthy fiber links.

The conversation also delves into the intricacies of the cluster itself. Rakesh outlines how Cisco's Silicon One architecture and Intelligent Collective Networking are engineered to handle synchronized surges of GPU traffic, mitigate bottlenecks, and expedite job completion time. Power efficiency is a crucial aspect of this discussion, encompassing the trade-off between network consumption and the energy allocated to GPUs.

Security and longevity are also addressed. The interview explores hardware-accelerated MACsec and IPsec, as well as the role of programmability in enabling infrastructure to adapt to the ever-evolving landscape of AI workloads. For infrastructure and datacenter leaders planning for more expansive, distributed AI environments, this discussion offers a pragmatic perspective on the networking hurdles that arise when a single site is no longer sufficient.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in AI

Theo Ported TypeScript to Rust with AI and Never Read the Code

Theo Browne published a Rust rewrite of the TypeScript compiler and admitted, in the README, that he has never read a line of the code. His own words: I've never read a line of this code.

  • Theo Browne created ts-rust, a Rust rewrite of TypeScript compiler without reading original code
  • $420,000 spent on API tokens for porting process, could have been $20,000
  • ts-rust 13x faster than tsc on VS Code, passes 181,711 Go tests at half the time

More from Wednesday 7 October →