Urgent.News

What's breaking now, across thousands of outlets.

AI

Nvidia’s dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents

Chipmaker Nvidia Corp. says its dedicated artificial intelligence inference accelerator Groq 3 LPX has now entered full production as it strives to maintain its dominance in the world of AI compute. The new chip, announced today at Hot Chips 2026, is described as a purpose-built extension to Nvidia’s flagship Vera Rubin data center platform. According […] The post Nvidia’s dedicated inference…

Nvidia’s dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents

Chipmaker Nvidia has announced that its dedicated AI inference accelerator, Groq 3 LPX, has entered full production. This purpose-built extension to Nvidia's flagship Vera Rubin data center platform is designed to deliver ultra-fast token generation speeds, essential for running responsive agentic AI workloads. Nebius Group N.V. is the first customer to commit to using the new chip.

Groq 3 LPX can offload decode workloads from the large-scale context processing to increase the speed at which AI agents can reason and work. The chipmaker claims it is four times more responsive for latency-sensitive workloads compared to rival platforms. This development marks another leap in AI throughput, efficiency, and responsiveness.

Nvidia's Grace Blackwell and NVL72 platforms have already revolutionized AI inference performance, and this new addition further strengthens its position in the AI compute market.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at siliconangle.com →

More in AI

How XPUs Meet a World-Class AI Factory

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators.

More from Monday 24 August →