Urgent.News

One page, thousands of outlets. See who else covered it.

Editions

AI

Nvidia wants to stop AI costs skyrocketing with its new software router — but will it really make a difference?

Nvidia says its new router cuts AI costs by up to 74%, but the partner that measured it says it couldn't prove the router beat using the cheaper model.

Nvidia wants to stop AI costs skyrocketing with its new software router — but will it really make a difference?

Nvidia has unveiled NeMo Switchyard, an open source model router designed to tackle soaring AI costs for enterprise customers. The software cleverly directs each agent request to the most cost-effective model capable of handling it, achieving a 74% reduction in costs compared to using only frontier models while only slightly compromising accuracy by 6%. This initiative joins a rising trend of players like RouteLLM, LiteLLM, and OpenRouter, as well as initiatives by OpenAI and AT&T, all aiming to cut costs in the AI space.

As AI adoption among enterprises continues to expand, so do costs, making Nvidia's solution particularly timely. The emerging AI routing field is proving lucrative, with major players like Stripe acquiring OpenRouter for over $7 billion. Nvidia's NeMo Switchyard functions much like a proxy, using a routing algorithm to decide which AI model to send queries to, while delivering answers in the required format for the calling application or user.

It supports OpenAI, Anthropic, and Responses API requests and provides transparency by documenting the selected model, decision rationale, token usage, and latency for each call.

One example cited by Nvidia is Nemotron Parse, a one-billion-parameter model designed for extracting structure from PDFs. While the routing approach may not always be feasible, it can lead to significant savings, especially when dealing with models like Nemotron 3.5 Lightning, where the judge model consumes up to 21.2% of the total cost.

However, there are tradeoffs, as routing costs can vary significantly, swinging between $2.16 and $3.61, depending on whether queries are escalated to more powerful models. This variability makes predicting costs challenging, although Nvidia maintains that routing helps lower average spend and pushes inference toward enterprise-owned hardware.

Written by urgent.news from TechRadar's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at techradar.com →

More in AI

Velaura AI raises $110M to develop power-efficient AI chips

Velaura AI Inc., a developer of low-power artificial intelligence chips, today announced that it has raised $110 million in funding. Seligman Ventures led the Series A round. It was joined by Samsung Catalyst Fund, Mayfield and more than a half-dozen others. Velaura is now valued at over $1 billion.

More from Wednesday 19 August →