{
  "id": 3863903,
  "title": "Nvidia and Cerebras are selling performance their customers will (probably) never see",
  "url": "https://urgent.news/2026/08/27/nvidia-and-cerebras-are-selling-performance-their-customers-will",
  "topic": "science",
  "section": "Science",
  "published": "2026-08-27T22:57:43.000Z",
  "source": {
    "name": "The Register Science",
    "slug": "the-register-science",
    "url": "https://www.theregister.com/systems/2026/08/27/nvidia-and-cerebras-are-selling-performance-their-customers-will-probably-never-see/5293117"
  },
  "original_language": "en",
  "account": "The Hot Chips conference in California saw Nvidia announce its new Groq-3-based LPX racks, with early tests showing the systems processing 3,400 tokens per second in Gemma 4 31B. This is four times faster than Cerebras' current offerings. However, Cerebras quickly responded by touting the performance of its next-gen CS-4 accelerators. While these top-line performance figures may feel instantaneous compared to current chatbots, they are more of a marketing gimmick rather than a realistic expectation for customers. The reality is that the economics of premium or ultra-low latency inference are more complex, and the true measure of performance lies in how efficiently the system can scale that speed to handle a large number of users. Both Nvidia and Cerebras are focusing on pushing the Pareto curve to the right, delivering massive bandwidth and extending performance, but the challenge lies in scaling that performance to accommodate a larger number of users. The key to making premium inference cost-effective lies in combining GPUs with Cerebras or Groq accelerators, as seen in partnerships between Nvidia, AMD, and AWS. While top-line performance makes for great headlines, the real benchmark should focus on achieving a balance between interactivity, concurrency, throughput, and economics.",
  "summary": "Touting batch 1 token generation is a bit like boasting about the top speed of your car",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 3,
    "also_reported_by": [
      {
        "outlet": "The Register",
        "title": "Nvidia and Cerebras are selling performance their customers will (probably) never see",
        "url": "https://urgent.news/2026/08/27/nvidia-and-cerebras-are-selling-performance-their-customers-will-3867811",
        "published": "2026-08-27T22:57:43.000Z"
      },
      {
        "outlet": "TechSpot",
        "title": "Modders got Nvidia's controversial DLSS 5 running early, and it's a huge performance hit",
        "url": "https://urgent.news/2026/08/28/modders-got-nvidias-controversial-dlss-5-running-early-and-its-a-huge",
        "published": "2026-08-28T02:41:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}