Cerebras CS-4 rack systems juice chips for every last drop of AI performance
Next-gen systems double per-chip performance while cramming 3x as many into a rack
Cerebras' latest wafer-scale accelerators, the CS-4 rack systems, are designed to maximize AI performance by addressing memory bandwidth limitations. The newly unveiled WSE-3T, or "Turbo," delivers twice the compute, memory fabric, and I/O bandwidth compared to its predecessor, the WSE-3, while maintaining the same process technology, wafer area, transistor count, core count, and SRAM capacity.
This improvement is achieved through enhanced power delivery, enabling higher operating frequencies and faster token generation. The WSE-3T boasts 250 petaFLOPS of AI compute, 44 GB of SRAM (43.2 PB/s memory bandwidth), and 2.4 Tbps of off-die connectivity, positioning it as a powerful contender in the AI landscape. However, Cerebras has strategically partnered with AWS and AMD to offload compute-intensive tasks, such as prompt processing, onto their respective Trainium XPUs and Instinct GPUs.
This approach leverages the SRAM-rich Cerebras chips for decoding tasks, reducing the need for numerous LPUs and enabling efficient inference. The modular, rack-scale design of the CS-4 rack systems allows for easier deployment, maintenance, and upgrades, with each system housing up to three backpacks containing accelerators. Although the increased power consumption is notable, it remains relatively conservative compared to upcoming 240 to 250 kW rack systems from AMD and Nvidia.
Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.