Nvidia and Cerebras are selling performance their customers will (probably) never see
Touting batch 1 token generation is a bit like boasting about the top speed of your car
The recent Hot Chips conference in California saw Nvidia boasting about its Groq-3-based LPX racks, which are said to be producing an impressive 3,400 tokens per second in Gemma 4 31B. This is four times faster than the rival Cerebras systems. However, Cerebras quickly responded by showcasing the performance of its next-gen CS-4 accelerators.
These figures, while impressive, may not be what customers will actually see in production. The companies are essentially debating numbers that may not be relevant for real-world use. While these numbers are akin to the top speed of a race car, they are not representative of the regular performance one might expect. The economics of premium or ultra-low latency inference are more complex and hinge on the ability to scale that performance efficiently.
The Pareto frontier chart illustrates the performance characteristics of various Nvidia B300 configurations across a range. While GPUs are great for high-throughput, low-interactivity applications, they struggle as the per-user generation rates increase. Companies like Nvidia and Cerebras are pushing the boundaries with their SRAM-heavy architectures, but they face limitations in terms of scalability.
These systems excel as decode accelerators, working best when combined with GPUs or other compute-heavy components. The sweet spot for performance often lies in the middle of the Pareto curve, balancing interactivity, concurrency, throughput, and economics.
Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.