{
  "id": 3065081,
  "title": "Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence (The Register)",
  "url": "https://urgent.news/2026/08/24/nvidia-says-its-groq-3-lpx-racks-delivered-3-400-tokens-per-second-in",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-24T16:30:00.000Z",
  "source": {
    "name": "Techmeme",
    "slug": "techmeme",
    "url": "https://www.theregister.com/systems/2026/08/24/what-nvidias-first-groq-3-lpu-benchmarks-do-and-dont-tell-us-about-its-20b-gamble/5291880"
  },
  "original_language": "en",
  "account": null,
  "summary": "Nvidia's Groq 3 LPX racks achieved 3,400 tokens per second in an Artificial Analysis benchmark. The test used Google's Gemma 4 31B model with a 100,000-token input sequence.\n\nAccording to Nvidia, this performance makes the Groq 3 LPX racks 4 times faster than the nearest alternative platform. The Register reports that this appears to be a reference to Cerebras, which managed 882 tokens per second under the same conditions.\n\nGroq's LPUs use an SRAM-heavy dataflow architecture designed for high-performance inference serving. This architecture relies on large pools of on-die SRAM, which provides faster memory bandwidth than traditional datacenter GPUs. The third generation of Groq's chips, launched as part of Nvidia's Vera Rubin platform, boasts 150 TB/s of memory bandwidth.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}