Urgent.News

What's breaking now, across thousands of outlets.

AI

Cisco Expands Secure AI Factory With NVIDIA

(RTTNews) - Cisco Systems, Inc. (CSCO), an American multinational technology conglomerate, announced on Tuesday that it is expanding its Secure AI Factory architecture with NVIDIA Corp. (NVDA), adding rack-scale computing solutions through a partnership with Supermicro to meet su

The AI infrastructure market is reaching a critical juncture, with the focus shifting from purchasing graphics processing units to constructing comprehensive AI factories capable of reliably, efficiently, and at scale generating tokens. These AI factories require a harmonious interplay between compute, networking, storage, cooling, software, and operations from the initial order to continuous production. If any component underperforms, valuable capacity remains dormant and revenue is lost.

In interviews with theCUBE, Cisco's Will Eatherton, senior vice president and head of networking engineering, and Nvidia's Senior Vice President of Networking Gilad Shainer, along with Marc Hamilton, vice president of solutions architecture and engineering, discussed the companies' strategies to address the execution challenge as neoclouds, sovereign AI programs, and enterprises deploy infrastructure into production.

They explored the necessary engineering, operational, and financial factors in building AI factories at rack scale, which is the impetus behind Cisco Secure AI Factory's expansion with Nvidia to encompass rack-scale systems.

By September, Cisco and Nvidia will offer customers the full rack-scale Secure AI Factory solution, which includes liquid-cooled compute, AI-optimized networking, validated designs, and unified operations. The objective is to expedite the transition from infrastructure ordering to the generation of the first token. The marketplace has entered the execution era, where the winners will be determined by the speed at which they convert capital expenditures into productive capacity.

Several factors contribute to this urgency. Neoclouds have customers waiting in line before GPUs arrive, implying that deployment delays directly result in deferred revenue. Similarly, enterprises face challenges as the cost of consuming models through external application programming interfaces increases. Sovereign AI programs add another layer of pressure, as they must strike a balance between performance, data control, and national infrastructure requirements.

The shift is economic. Traditional enterprise infrastructure was often seen as a cost center to be optimized downward. However, an AI factory is designed to produce digital products - tokens. These tokens power applications, agents, and business processes, providing infrastructure with a direct link to revenue. Consequently, utilization, availability, tokens per second, and tokens per watt have become crucial business metrics, rather than mere engineering measurements.

Marc Hamilton emphasized that the real cost savings in an AI factory do not stem from cost reduction but from driving revenue - ensuring a repeatable method to achieve this. Building an AI factory is not about connecting components and hoping for the best. Instead, it necessitates constructing a supercomputer that must be assembled rapidly, delivering the highest number of tokens per second, power efficiency, and other critical metrics.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at nasdaq.com →

More in AI

Cisco and Nvidia take AI factories from rack to runtime

AI factories are moving from ambitious plans toward production, but the path from graphics processing unit acquisition to usable systems remains a race against time. Neoclouds already have customers waiting for capacity, enterprises are looking to bring inference workloads closer to home and sovereign AI programs are being built now.

More from Tuesday 25 August →