Urgent.News

What's breaking now, across thousands of outlets.

Tech

What Nvidia's first Groq 3 LPU benchmarks do and don't tell us about its $20B gamble

Gemma 4 31B performance tests offer a best-case scenario for next-gen dataflow accelerators

What Nvidia's first Groq 3 LPU benchmarks do and don't tell us about its $20B gamble

Nvidia's $20 billion investment in Groq's LPU technology appears to have paid off, as evidenced by a recent benchmark conducted by Artificial Analysis. The benchmark revealed that Nvidia's LPX rack systems, powered by Groq 3-based LPUs, achieved a token-generation rate of 3,400 tokens per second (tok/s) when processing a 100,000-token input sequence from Google's Gemma 4 31B model.

This performance is reported to be four times faster than the closest alternative platform, Cerebras, which delivered a respectable 882 tok/s under similar conditions.

Groq's LPUs employ an SRAM-heavy dataflow architecture that excels at high-performance inference serving. Unlike conventional datacenter GPUs that rely on high-speed DRAM memory (GDDR7 and HBM4), Groq's chips are entirely powered by on-die SRAM, which boasts a staggering 2.75 TB/s bandwidth, far surpassing even the best HBM stacks available today.

Despite the SRAM's efficiency, its limited capacity is a notable drawback. For instance, Groq 3 LPU chips only provide 500 MB of SRAM, a fraction of the onboard memory (288 GB) found in Nvidia's top-spec Rubin GPU. To circumvent this limitation, Nvidia's architecture utilizes Ethernet to distribute models across multiple accelerators, enabling the LPX rack to support up to 256 LPUs, yielding 128 GB of high-bandwidth SRAM.

The significance of these benchmark results lies in the potential impact on AI agents and agentic systems. The faster inference servers can generate tokens, the more complex a model can reason, and the more turns the agent can take in a given timeframe. Nvidia's LPX system, in combination with Groq 3 LPUs, has evidently captured the attention of Nvidia's customers.

Just recently, Netherlands-based neocloud Nebius announced its intention to deploy the combined systems in its datacenters. However, it remains to be seen how well the architecture will scale to larger, more intricate models.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at theregister.com →

More in Tech

More from Monday 24 August →