{
  "id": 3079762,
  "title": "What Nvidia's first Groq 3 LPU benchmarks tell us about its $20B gamble",
  "url": "https://urgent.news/2026/08/24/what-nvidias-first-groq-3-lpu-benchmarks-tell-us-about-its-20b-gamble-3079762",
  "topic": "science",
  "section": "Science",
  "published": "2026-08-24T15:00:00.000Z",
  "source": {
    "name": "The Register Science",
    "slug": "the-register-science",
    "url": "https://www.theregister.com/systems/2026/08/24/what-nvidias-first-groq-3-lpu-benchmarks-tell-us-about-its-20b-gamble/5291880"
  },
  "original_language": "en",
  "account": "Nvidia's $20 billion investment in Groq's Low-Power Unit (LPU) technology has proven successful, according to the first benchmarks released by the company. Nvidia's LPX rack systems, powered by Groq 3 LPUs, achieved a remarkable 3,400 tokens per second (tok/s) when processing a 100,000-token input sequence using Google's Gemma 4 31B model. This represents a fourfold speedup compared to the nearest alternative platform, which Artificial Analysis' leaderboard lists at 882 tok/s. Groq's LPUs employ a dataflow architecture with SRAM (Static Random Access Memory), which is significantly faster than traditional datacenter GPUs that rely on GDDR7 and HBM4 memory technologies. Although each Groq 3 LPU has only 500 MB of memory, far less than the 288 GB onboard in Nvidia's top-specced Rubin GPU, the architecture compensates by distributing models across multiple accelerators using Ethernet connectivity. This allows LPX racks to be equipped with up to 256 LPUs, providing 128 GB of high-bandwidth SRAM capacity. The benefits of faster inference servers become apparent in agentic AI applications, where the speed of generating tokens directly impacts the reasoning capabilities of AI agents and their ability to process information and take actions. While the results are impressive, Nvidia acknowledges that the model used in the benchmark is relatively small and dense, with 31 billion parameters all activated in each token generation. Scaling this architecture to larger, more complex models would require numerous LPUs, making it a significant undertaking. Despite the challenges, Nvidia's combination of GPUs and Groq 3 LPUs appears to have garnered significant attention from customers, with Netherlands-based neocloud Nebius among the first to deploy the combined systems in its datacenters.",
  "summary": "Gemma 4 31B performance tests offer a best-case scenario for next-gen dataflow accelerators",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "The Register",
        "title": "What Nvidia's first Groq 3 LPU benchmarks tell us about its $20B gamble",
        "url": "https://urgent.news/2026/08/24/what-nvidias-first-groq-3-lpu-benchmarks-tell-us-about-its-20b-gamble",
        "published": "2026-08-24T15:00:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}