Cerebras CS-4 rack systems juice chips for every last drop of AI performance
Next-gen systems double per-chip performance while cramming 3x as many into a rack
Cerebras has unveiled its next-generation Wafer Scale Engine (WSE) and Nexus rack systems, aiming to further extend its lead in AI performance. The WSE-3T, or "Turbo," promises twice the compute, memory fabric, and I/O bandwidth of the previous WSE-3, all while using the same process technology and wafer area size. This is achieved through improved power delivery, which enables higher operating frequencies and faster token generation.
The WSE-3T boasts 250 petaFLOPS of AI compute, 44 GB of SRAM, 43.2 PB/s of memory bandwidth, and 2.4 Tbps of off-die connectivity.
Cerebras has also partnered with Amazon Web Services (AWS) and AMD to offload compute-intensive prompt processing bits of the inference pipeline, treating its chips as decode accelerators rather than primary AI processors. This approach is similar to Nvidia's use of Groq LPUs in its LPX rack systems. However, Cerebras opted to double performance this generation instead of increasing SRAM capacity, which hasn't seen a significant boost since the WSE-2 launch five years ago.
With the new rack-scale compute architecture, Cerebras has moved from a monolithic system to a modular design, featuring "backpack" form factors and dedicated power shelves. Each CS-4 can be equipped with up to three backpacks, housing the accelerators, and the front of the rack is dedicated to power delivery. The company claims that their more efficient power delivery allows pushing twice as many watts through the chip, resulting in a system power range of 120 kW to 140 kW for the CS-4 rack.
Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.