{
  "id": 3867811,
  "title": "Nvidia and Cerebras are selling performance their customers will (probably) never see",
  "url": "https://urgent.news/2026/08/27/nvidia-and-cerebras-are-selling-performance-their-customers-will-3867811",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-08-27T22:57:43.000Z",
  "source": {
    "name": "The Register",
    "slug": "the-register",
    "url": "https://www.theregister.com/systems/2026/08/27/nvidia-and-cerebras-are-selling-performance-their-customers-will-probably-never-see/5293117"
  },
  "original_language": "en",
  "account": "The recent Hot Chips conference in California saw Nvidia boasting about its Groq-3-based LPX racks, which are said to be producing an impressive 3,400 tokens per second in Gemma 4 31B. This is four times faster than the rival Cerebras systems. However, Cerebras quickly responded by showcasing the performance of its next-gen CS-4 accelerators. These figures, while impressive, may not be what customers will actually see in production. The companies are essentially debating numbers that may not be relevant for real-world use. While these numbers are akin to the top speed of a race car, they are not representative of the regular performance one might expect. The economics of premium or ultra-low latency inference are more complex and hinge on the ability to scale that performance efficiently. The Pareto frontier chart illustrates the performance characteristics of various Nvidia B300 configurations across a range. While GPUs are great for high-throughput, low-interactivity applications, they struggle as the per-user generation rates increase. Companies like Nvidia and Cerebras are pushing the boundaries with their SRAM-heavy architectures, but they face limitations in terms of scalability. These systems excel as decode accelerators, working best when combined with GPUs or other compute-heavy components. The sweet spot for performance often lies in the middle of the Pareto curve, balancing interactivity, concurrency, throughput, and economics.",
  "summary": "Touting batch 1 token generation is a bit like boasting about the top speed of your car",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 3,
    "also_reported_by": [
      {
        "outlet": "The Register Science",
        "title": "Nvidia and Cerebras are selling performance their customers will (probably) never see",
        "url": "https://urgent.news/2026/08/27/nvidia-and-cerebras-are-selling-performance-their-customers-will",
        "published": "2026-08-27T22:57:43.000Z"
      },
      {
        "outlet": "TechSpot",
        "title": "Modders got Nvidia's controversial DLSS 5 running early, and it's a huge performance hit",
        "url": "https://urgent.news/2026/08/28/modders-got-nvidias-controversial-dlss-5-running-early-and-its-a-huge",
        "published": "2026-08-28T02:41:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}