Urgent.News

What's breaking now, across thousands of outlets.

Tech

Nvidia says Groq 3 LPX rack enters full production

Nvidia said each rack contains 256 Groq 3 chips manufactured by Samsung.

Nvidia says Groq 3 LPX rack enters full production

Nvidia's investment of $20 billion in Groq's Low-Power Unit (LPU) technology appears to have paid off, as demonstrated by the first benchmark results from Nvidia's LPX rack systems. The independent benchmark conducted by Artificial Analysis revealed that Nvidia's LPX rack systems can process 3,400 tokens per second (tok/s) with a 100,000-token input sequence in Google's Gemma 4 31B model.

This performance is reportedly four times faster than the nearest alternative platform, which could be a direct reference to Cerebras, which achieved 882 tok/s under similar conditions.

Groq's LPUs feature a dataflow architecture centered around SRAM, which is significantly faster than the high-speed DRAM memory technology (GDDR7 and HBM4) used by traditional datacenter GPUs. SRAM boasts around 2.75 TB/s of bandwidth, making it a bottleneck to overcome for inference tasks. The third-generation LPUs in Nvidia's broader Vera Rubin platform offer 150 TB/s of memory bandwidth.

However, the compact nature of SRAM limits the onboard memory to just 500 MB per Groq 3 LPU, which is insufficient to run the 31B Gemma model on a single LPU.

To address this limitation, Nvidia employs Ethernet to distribute models across multiple accelerators. Each LPX rack can accommodate up to 256 LPUs, providing 128 GB of high-bandwidth SRAM. For larger models, multiple LPX racks can be interconnected. The purpose behind running the 31B Gemma model at 3,400 tok/s is unclear, as it is a relatively small model. However, the faster inference could benefit AI code assistants or agents, allowing them to process more information and take actions more quickly.

The impressive performance of the LPX system has caught Nvidia's customers' attention, as evidenced by the inclusion of the Netherlands-based neocloud Nebius among the first users of the combined Nvidia GPUs and Groq 3 LPUs in their datacenters. Despite the impressive tok/s figure, the benchmark results are still subject to limitations, such as the dense model's extensive computational requirements.

Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at techinasia.com →

More in Tech

[Sponsor] Finalist 4: Inspired by Paper Day Planners

Finalist 4 is the biggest update yet to the paper-inspired day planner for iPhone, iPad, Mac, and now Apple Watch. The headline is Notes, in the app or as a folder of Markdown files that round-trips…

  • Finalist releases version 4.0 with notes, watch app, and daily page
  • Users can drag, resize, and accept tasks without affecting real tasks
  • App aims to simplify productivity by integrating tasks and thoughts

I built an open cross-reference dataset for 217 global steel grades so engineers stop guessing "is that equivalent?"

__ If you've ever received a spec that says ASTM A36 while your supplier quotes S275JR , you know the moment of dread: are these actually equivalent, or am I about to spec the wrong material into a…

  • Dataset of 217 steel grades created to eliminate guessing in engineering
  • Includes normalized properties, international equivalents, and physical characteristics
  • Material comparison engine enables side-by-side analysis of mechanical properties

I built a free image and video hosting tool after Imgur blocked the UK

On 30 September 2025, Imgur blocked the entire United Kingdom. No warning. No migration tool. No grace period. One day it worked, the next it didn't — and with it went millions of embedded images…

  • Imgur blocked UK without notice on Sep 30, 2025
  • DBimg offers free, no-account image/video hosting
  • Service supports major formats, CDN for fast delivery

More from Tuesday 25 August →