NVIDIA has introduced Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model designed for high-volume tasks within long-running agentic AI systems.
NVIDIA says the model delivers up to four times faster output and 30% faster agentic task completion than comparable models. It can run across RTX PCs, DGX systems, Jetson devices, data centers, and cloud environments.
NVIDIA also launched NeMo Switchyard, an open-source routing library that automatically directs requests to suitable models based on quality, latency, and cost requirements. Internal benchmarks show Switchyard can maintain frontier-level accuracy while reducing task completion costs to nearly one-third of using Opus 4.8 alone.




