Nvidia launches a smaller, faster Nemotron model and a router to put it to work
Nvidia on Tuesday launched Nemotron 3.5 Lightning, the newest member of its Nemotron 3 family of open models. In addition, The post Nvidia launches a smaller, faster Nemotron model and a router to put it to work appeared first on The New Stack .
Nvidia has introduced Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, and NeMo Switchyard, an open-source library for routing models. Nemotron 3.5 Lightning, developed with contributions from the Nemotron coalition, exhibits reasoning capabilities similar to the larger Nemotron 3 Super model, but falls short of Google's Gemma 4 31B model in terms of performance.
Nvidia emphasizes speed and customization, claiming that 3.5 Lightning can deliver up to 4x faster output speeds and can be finely tuned for specific workflows. Post-training customization significantly improves accuracy and outperforms proprietary models in specialized tasks. NeMo Switchyard, a new routing library, allows developers to define model pools and set routing criteria based on quality, latency, or cost.
In Nvidia's internal benchmarks, a Switchyard-routed system using a combination of open models and Anthropic's Opus 4.8 maintained frontier-level accuracy while cutting task-completion costs to about a third of running Opus alone. The Nemotron 3.5 Lightning model is available on Hugging Face, ModelScope, OpenRouter, and Nvidia's NIM microservice platform, while NeMo Switchyard is open-source and available on GitHub, with more partner integrations planned.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — it may contain errors, so check the original before relying on it.