NVIDIA Nemotron 3.5 Lightning now available in Amazon SageMaker JumpStart
NVIDIA Nemotron 3.5 Lightning, an open model built for high-volume agentic workloads, is now available in Amazon SageMaker JumpStart. This post shows how to deploy the 30B Mixture-of-Experts model (3B active), which delivers up to 4x higher throughput and up to 30% faster task completion for always-on agents.
NVIDIA Nemotron 3.5 Lightning, a specialized model for high-volume agentic workloads, is now available on Amazon SageMaker JumpStart. This open model, developed from NVIDIA’s Nemotron 3 Ultra, is designed for always-on agents that require specialized model execution. With up to 30B total parameters but only 3B active, it can run efficiently on a single GPU and is capable of delivering up to 4x higher throughput and up to 30% faster task completion.
It supports a context length of up to 1M tokens and utilizes a Hybrid Mixture-of-Experts (MoE) architecture. The model can be deployed without the need for configuring serving infrastructure, and it allows users to customize it for domain-specific accuracy. It is particularly suited for personal assistants, financial services, cybersecurity operations, telecom, and retail applications, among others.
Written by urgent.news from AWS Machine Learning's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.