AWS has made NVIDIA Nemotron 3.5 Lightning available through Amazon SageMaker JumpStart, simplifying deployment for high-volume agentic AI workloads. The open model uses a hybrid mixture-of-experts architecture with 30 billion total parameters and 3 billion active parameters, allowing it to run on a single supported GPU.
It supports context windows up to 1 million tokens and DFlash speculative decoding. NVIDIA reports up to four times higher throughput and 30% faster task completion for specialized agent workloads.
Developers can deploy BF16 and NVFP4 variants through SageMaker JumpStart, Hugging Face, or the SageMaker Python SDK without configuring serving infrastructure manually.




