What Nvidia's first Groq 3 LPU benchmarks tell us about its $20B gamble
Gemma 4 31B performance tests offer a best-case scenario for next-gen dataflow accelerators
Nvidia's $20 billion investment in Groq's Low-Power Unit (LPU) technology has yielded impressive results, according to the first set of benchmarks. In an independent test conducted by Artificial Analysis, Nvidia's LPX rack systems demonstrated a token generation speed of 3,400 tokens per second (tok/s) with a 100,000-token input sequence from Google's Gemma 4 31B model.
This performance is four times faster than the nearest competitor, Cerebras, which achieved 882 tok/s in the same conditions. Groq's LPUs utilize a dataflow architecture with a high-speed SRAM design, which is significantly faster than traditional datacenter GPUs that rely on DRAM memory technology. While each LPU has only 500 MB of memory, it is distributed across multiple LPUs using Ethernet to handle larger models.
This architecture allows Nvidia to achieve high-speed inference while maintaining a relatively compact design. The company is targeting the agentic AI market, where faster inference servers can enhance the capabilities of AI agents and justify a premium price. However, the success of this approach will depend on how effectively Nvidia can scale the technology to handle larger, more complex models.
The benchmark results suggest that Nvidia's LPX system is already impressive, but there is still uncertainty about its long-term performance and scalability.
Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.