What Nvidia's first Groq 3 LPU benchmarks do and don't tell us about its $20B gamble
Gemma 4 31B performance tests offer a best-case scenario for next-gen dataflow accelerators
Nvidia's investment of $20 billion in Groq's Low-Power Unit (LPU) technology appears to have paid off, as demonstrated by the first benchmark results from Nvidia's LPX rack systems. The independent benchmark conducted by Artificial Analysis revealed that Nvidia's LPX rack systems can process 3,400 tokens per second (tok/s) with a 100,000-token input sequence in Google's Gemma 4 31B model.
This performance is reportedly four times faster than the nearest alternative platform, which could be a direct reference to Cerebras, which achieved 882 tok/s under similar conditions.
Groq's LPUs feature a dataflow architecture centered around SRAM, which is significantly faster than the high-speed DRAM memory technology (GDDR7 and HBM4) used by traditional datacenter GPUs. SRAM boasts around 2.75 TB/s of bandwidth, making it a bottleneck to overcome for inference tasks. The third-generation LPUs in Nvidia's broader Vera Rubin platform offer 150 TB/s of memory bandwidth.
However, the compact nature of SRAM limits the onboard memory to just 500 MB per Groq 3 LPU, which is insufficient to run the 31B Gemma model on a single LPU.
To address this limitation, Nvidia employs Ethernet to distribute models across multiple accelerators. Each LPX rack can accommodate up to 256 LPUs, providing 128 GB of high-bandwidth SRAM. For larger models, multiple LPX racks can be interconnected. The purpose behind running the 31B Gemma model at 3,400 tok/s is unclear, as it is a relatively small model. However, the faster inference could benefit AI code assistants or agents, allowing them to process more information and take actions more quickly.
The impressive performance of the LPX system has caught Nvidia's customers' attention, as evidenced by the inclusion of the Netherlands-based neocloud Nebius among the first users of the combined Nvidia GPUs and Groq 3 LPUs in their datacenters. Despite the impressive tok/s figure, the benchmark results are still subject to limitations, such as the dense model's extensive computational requirements.
Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.