Urgent.News

What's breaking now, across thousands of outlets.

AI

AI factories enter the execution era as Cisco and NVIDIA push rack-scale systems into production

The artificial intelligence infrastructure market is crossing an important threshold. The conversation is shifting from acquiring graphics processing units to building complete AI factories that can generate tokens reliably, efficiently and at scale. This is the next bottleneck. GPUs may be the engine, but an AI factory is a system. Compute, networking, storage, cooling, software […] The post AI…

AI factories enter the execution era as Cisco and NVIDIA push rack-scale systems into production

The AI infrastructure market is reaching a critical juncture, with the focus shifting from purchasing graphics processing units to constructing comprehensive AI factories capable of reliably, efficiently, and at scale generating tokens. These AI factories require a harmonious interplay between compute, networking, storage, cooling, software, and operations from the initial order to continuous production. If any component underperforms, valuable capacity remains dormant and revenue is lost.

In interviews with theCUBE, Cisco's Will Eatherton, senior vice president and head of networking engineering, and Nvidia's Senior Vice President of Networking Gilad Shainer, along with Marc Hamilton, vice president of solutions architecture and engineering, discussed the companies' strategies to address the execution challenge as neoclouds, sovereign AI programs, and enterprises deploy infrastructure into production.

They explored the necessary engineering, operational, and financial factors in building AI factories at rack scale, which is the impetus behind Cisco Secure AI Factory's expansion with Nvidia to encompass rack-scale systems.

By September, Cisco and Nvidia will offer customers the full rack-scale Secure AI Factory solution, which includes liquid-cooled compute, AI-optimized networking, validated designs, and unified operations. The objective is to expedite the transition from infrastructure ordering to the generation of the first token. The marketplace has entered the execution era, where the winners will be determined by the speed at which they convert capital expenditures into productive capacity.

Several factors contribute to this urgency. Neoclouds have customers waiting in line before GPUs arrive, implying that deployment delays directly result in deferred revenue. Similarly, enterprises face challenges as the cost of consuming models through external application programming interfaces increases. Sovereign AI programs add another layer of pressure, as they must strike a balance between performance, data control, and national infrastructure requirements.

The shift is economic. Traditional enterprise infrastructure was often seen as a cost center to be optimized downward. However, an AI factory is designed to produce digital products - tokens. These tokens power applications, agents, and business processes, providing infrastructure with a direct link to revenue. Consequently, utilization, availability, tokens per second, and tokens per watt have become crucial business metrics, rather than mere engineering measurements.

Marc Hamilton emphasized that the real cost savings in an AI factory do not stem from cost reduction but from driving revenue - ensuring a repeatable method to achieve this. Building an AI factory is not about connecting components and hoping for the best. Instead, it necessitates constructing a supercomputer that must be assembled rapidly, delivering the highest number of tokens per second, power efficiency, and other critical metrics.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at siliconangle.com →

More in AI

More from Tuesday 25 August →