Urgent.News

What's breaking now, across thousands of outlets.

AI

Nvidia's latest solution to soaring enterprise AI costs is...a router?

NeMo Switchyard brings GPT-5-style model routing to the mainstream

Nvidia's latest solution to soaring enterprise AI costs is...a router?

Nvidia's latest solution to high enterprise AI costs is a router called Switchyard. Introduced alongside the Nemotron 3.5-30B-A3B-Lightning model, Switchyard aims to optimize AI spending by intelligently routing prompts to different models based on factors like cost, latency, or output quality. By directing some requests to smaller, cheaper models, Nvidia claims Switchyard can slash job completion costs by 74 percent compared to using Claude Opus 4.8, albeit with a six-point accuracy tradeoff.

The router simplifies the process of choosing the right model for specific tasks, as enterprises often struggle with determining when to use larger, more powerful models and when to opt for smaller, more specialized alternatives. Nvidia has spent years developing task-specific models and has introduced several app-specific options, like Nemotron Parse, which excels at PDF analysis.

By offloading simpler tasks to specialized models, enterprises can improve accuracy while reducing costs. While the concept of model routers isn't new, Nvidia's Switchyard abstracts the process, making it more accessible for enterprise AI adoption.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in AI

More from Wednesday 12 August →