Nvidia says Groq racks will be online this year following $20 billion purchase
Nvidia's race to manufacture Groq chips and make them available to customers highlights the growing importance in AI of low-latency inference.
Chipmaker Nvidia has announced that its dedicated AI inference accelerator, Groq 3 LPX, has entered full production. This purpose-built extension to Nvidia's flagship Vera Rubin data center platform is designed to deliver ultra-fast token generation speeds, essential for running responsive agentic AI workloads. Nebius Group N.V. is the first customer to commit to using the new chip.
Groq 3 LPX can offload decode workloads from the large-scale context processing to increase the speed at which AI agents can reason and work. The chipmaker claims it is four times more responsive for latency-sensitive workloads compared to rival platforms. This development marks another leap in AI throughput, efficiency, and responsiveness.
Nvidia's Grace Blackwell and NVL72 platforms have already revolutionized AI inference performance, and this new addition further strengthens its position in the AI compute market.
Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.