Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence (The Register)
Nvidia's $20 billion bet on Groq's LPU tech sure looks like it was a good one. On Monday, the GPU giant offered the first glimpse …
Nvidia's Groq 3 LPX racks achieved 3,400 tokens per second in an Artificial Analysis benchmark. The test used Google's Gemma 4 31B model with a 100,000-token input sequence.
According to Nvidia, this performance makes the Groq 3 LPX racks 4 times faster than the nearest alternative platform. The Register reports that this appears to be a reference to Cerebras, which managed 882 tokens per second under the same conditions.
Groq's LPUs use an SRAM-heavy dataflow architecture designed for high-performance inference serving. This architecture relies on large pools of on-die SRAM, which provides faster memory bandwidth than traditional datacenter GPUs. The third generation of Groq's chips, launched as part of Nvidia's Vera Rubin platform, boasts 150 TB/s of memory bandwidth.
Brief written by urgent.news from Techmeme, The Register — 2 reports on this story. Machine-written — may contain errors; check the original before relying on it.