NVIDIA has introduced Nemotron 3 Super, a 120B-parameter open large language model designed specifically for high-throughput, agentic AI systems such as collaborative agents, tool-using assistants, and IT automation workflows. The model uses a hybrid Latent Mixture-of-Experts (LatentMoE) architecture that interleaves Mamba-2, MoE, and attention layers, activating only about 12B parameters per token to deliver significantly higher inference efficiency than dense models of similar scale. NVIDIA positions Nemotron 3 Super as part of its Nemotron 3 family (Nano, Super, Ultra), focusing this tier on multi-agent coordination, long-context reasoning, and high-volume enterprise workloads. Nemotron 3 Super offers a context window of up to 1 million tokens, aimed at long-horizon planning, cross-document reasoning, and multi-user or multi-agent conversations. NVIDIA reports that Super achieves up to ~5x higher throughput than its previous Nemotron models and more than 50% higher token generation speed than leading open peers, helped by multi-token prediction (MTP) and the LatentMoE design, which can consult multiple experts at the cost of one. The model is trained with NVFP4 quantization for compute efficiency and incorporates multi-environment reinforcement learning across more than ten environments to improve reasoning and tool-use performance, with strong results on benchmarks such as AIME 2025, TerminalBench, and SWE-Bench Verified. Strategically, NVIDIA is releasing Nemotron 3 Super as an open model with downloadable weights, data recipes, and a permissive commercial license, under the NVIDIA Nemotron Open Model License, allowing self-hosting, fine-tuning, and on-prem deployment. The model is distributed via platforms such as NVIDIA NIM, major cloud providers, and hosting services including Hugging Face, OpenRouter, and others, and can be run on multi-GPU server configurations (e.g., 8× H100) or, in more optimized/distilled forms, on smaller setups with around 64GB memory for local experimentation. For the broader AI ecosystem, Nemotron 3 Super underscores NVIDIA’s push to pair its GPU and inference stack with high-performance open models optimized for agentic AI, giving enterprises and developers an alternative to closed proprietary models for long-context reasoning, automation, and multi-agent systems.

AI-generated background, compiled from web sources — not editorial content.

More coverage

Explore the topic

More on AI

Comments