NVIDIA has detailed support for Meta’s Muse Glimmer, a 30-billion-parameter open-weight dense model designed for local, long-running agentic AI workloads.
Featuring a 120K-plus context window, the model activates every parameter per token to provide predictable latency, reliable instruction following, and sustained long-context performance.
Muse Glimmer can run fully on-device across NVIDIA platforms including GeForce RTX 5090, DGX Spark, DGX Station, and Jetson, helping keep sensitive data local. NVIDIA reports throughput exceeding 20,000 tokens per second per GPU on Blackwell Ultra and supports deployment through NVIDIA NIM, SGLang, vLLM, and NemoClaw for agentic workflows.





