Urgent.News

What's breaking now, across thousands of outlets.

AI

Nvidia's latest solution to soaring enterprise AI costs is...a router?

NeMo Switchyard brings GPT-5-style model routing to the mainstream

Nvidia's latest solution to soaring enterprise AI costs is...a router?

Nvidia's new software platform, NeMo Switchyard, is designed to address enterprise AI costs by acting as a router that can direct AI model requests based on factors such as cost, latency, and output quality. By routing certain prompts to smaller, cheaper models, Nvidia claims Switchyard can reduce job completion costs by up to 74 percent compared to using a single, larger model.

However, this cost savings comes at the expense of approximately six points of accuracy. Switchyard is part of Nvidia's broader strategy to make enterprise AI spend more manageable by offering a mix of high-performance open weights models like Nemotron 3.5-30B-A3B-Lightning and application-specific models tailored for specific tasks, such as Nemotron Parse, which excels at extracting context from PDF documents.

While the concept of model routers is not new, with OpenAI employing a similar system for GPT-5 and AT&T using a "smart router" for certain applications, Nvidia's approach aims to simplify the process of routing requests to the most appropriate model for the job. The ultimate goal is to make enterprise AI adoption more practical and cost-effective, benefiting Nvidia's bottom line.

Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at theregister.com →

More in AI

More from Wednesday 12 August →