{
  "id": 3065496,
  "title": "What Nvidia's first Groq 3 LPU benchmarks do and don't tell us about its $20B gamble",
  "url": "https://urgent.news/2026/08/24/what-nvidias-first-groq-3-lpu-benchmarks-do-and-dont-tell-us-about-3065496",
  "topic": "science",
  "section": "Science",
  "published": "2026-08-24T15:00:00.000Z",
  "source": {
    "name": "The Register Science",
    "slug": "the-register-science",
    "url": "https://www.theregister.com/systems/2026/08/24/what-nvidias-first-groq-3-lpu-benchmarks-do-and-dont-tell-us-about-its-20b-gamble/5291880"
  },
  "original_language": "en",
  "account": "Nvidia's investment of $20 billion in Groq's Low-Power Unit (LPU) technology appears to have paid off, as demonstrated by the first benchmark results from Nvidia's LPX rack systems. The independent benchmark conducted by Artificial Analysis revealed that Nvidia's LPX rack systems can process 3,400 tokens per second (tok/s) with a 100,000-token input sequence in Google's Gemma 4 31B model. This performance is reportedly four times faster than the nearest alternative platform, which could be a direct reference to Cerebras, which achieved 882 tok/s under similar conditions.\n\nGroq's LPUs feature a dataflow architecture centered around SRAM, which is significantly faster than the high-speed DRAM memory technology (GDDR7 and HBM4) used by traditional datacenter GPUs. SRAM boasts around 2.75 TB/s of bandwidth, making it a bottleneck to overcome for inference tasks. The third-generation LPUs in Nvidia's broader Vera Rubin platform offer 150 TB/s of memory bandwidth. However, the compact nature of SRAM limits the onboard memory to just 500 MB per Groq 3 LPU, which is insufficient to run the 31B Gemma model on a single LPU.\n\nTo address this limitation, Nvidia employs Ethernet to distribute models across multiple accelerators. Each LPX rack can accommodate up to 256 LPUs, providing 128 GB of high-bandwidth SRAM. For larger models, multiple LPX racks can be interconnected. The purpose behind running the 31B Gemma model at 3,400 tok/s is unclear, as it is a relatively small model. However, the faster inference could benefit AI code assistants or agents, allowing them to process more information and take actions more quickly.\n\nThe impressive performance of the LPX system has caught Nvidia's customers' attention, as evidenced by the inclusion of the Netherlands-based neocloud Nebius among the first users of the combined Nvidia GPUs and Groq 3 LPUs in their datacenters. Despite the impressive tok/s figure, the benchmark results are still subject to limitations, such as the dense model's extensive computational requirements.",
  "summary": "Gemma 4 31B performance tests offer a best-case scenario for next-gen dataflow accelerators",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "The Register",
        "title": "What Nvidia's first Groq 3 LPU benchmarks do and don't tell us about its $20B gamble",
        "url": "https://urgent.news/2026/08/24/what-nvidias-first-groq-3-lpu-benchmarks-do-and-dont-tell-us-about",
        "published": "2026-08-24T15:00:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}