Urgent.News

What's breaking now, across thousands of outlets.

AI

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

The next era of AI inference won’t be defined by a single breakthrough chip, network or system. It’ll be defined by how every layer of the AI factory works together. That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems. Announced today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq […]

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

NVIDIA's NVIDIA Groq 3 LPX, a low-latency inference accelerator, is now in full production and designed to enhance NVIDIA's Vera Rubin NVL72 AI factory platform. The new accelerator is specifically engineered to accelerate token generation for agentic systems, which are increasingly generating more tokens, processing larger context windows, and collaborating with other AI systems to solve complex problems.

The Groq 3 LPX achieves 3,400 output tokens per second for 100,000-token long-context use cases, offering a significant performance boost over the nearest alternative platform. Industry leaders, such as Nebius, SpaceXAI, and CoreWeave, are adopting the Vera Rubin platform to power their agentic AI applications. CoreWeave is deploying Spectrum-X Multiplane, a high-bandwidth, flat and lossless AI network solution, to connect NVIDIA Vera Rubin racks, unlocking new advances for its AI cloud infrastructure.

Meanwhile, SpaceXAI plans to build and scale its future AI architecture around NVIDIA Vera CPUs, integrating them with NVIDIA accelerated computing, networking, and software to advance AI at an unprecedented scale. NVIDIA's extreme codesign approach integrates compute, networking, and inference acceleration into a unified system, optimizing every stage of the AI pipeline from context and communication to generation.

This integrated AI factory architecture is designed to deliver low latency, extreme throughput, and scalable economics, turning AI factories into integrated engines for intelligence capable of turning ever-growing volumes of tokens into revenue.

Written by urgent.news from NVIDIA Blog's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at blogs.nvidia.com →

More in AI

More from Monday 24 August →