Urgent.News

What's breaking now, across thousands of outlets.

AI

AI infrastructure reliability moves beyond the rack

AI infrastructure reliability depends on how well compute, networking, power and cooling operate together. Turning early hardware breakthroughs into repeatable deployments requires coordination across engineering, manufacturing and operations. Dell Technologies Inc.’s partnership with CoreWeave Inc. has helped it adapt successive generations of rack-scale systems to different facilities,…

AI infrastructure reliability moves beyond the rack

AI infrastructure reliability is becoming increasingly complex as it moves beyond the rack. Dell Technologies' partnership with CoreWeave has played a crucial role in adapting rack-scale systems to various facilities. Sarat Krishnan, director of PowerEdge AI architecture and systems development engineering at Dell, explained how cooling and power requirements differ depending on the data center's design.

They have learned to create modular rack-scale infrastructure that can quickly adapt to different cooling and power needs.

CoreWeave's Racky rack manager integrates power, cooling, and environmental telemetry into a unified control interface. This integration allows Dell to refine diagnostics across server and rack assembly using operating feedback from CoreWeave. By shifting complex diagnostics left in the manufacturing process, they can mitigate risks and prevent costly issues later.

As AI infrastructure evolves, it is no longer solely confined to the rack. Instead, entire data halls or data centers are now considered the "computer." This shift in perspective requires tight systems integration across the entire stack, encompassing compute trays, switches, data processing units, and network fabrics.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at siliconangle.com →

More in AI

More from Monday 5 October →