NVIDIA launches Nemotron 3 Super, a 120B-parameter open model designed for agentic AI systems with up to 5x higher throughput


12 recorded changes
Want your article here?
Promote with Leviathan News

12 recorded changes
Want your article here?
Promote with Leviathan NewsNVIDIA has introduced Nemotron 3 Super, a 120B-parameter open large language model designed specifically for high-throughput, agentic AI systems such as collaborative agents, tool-using assistants, and IT automation workflows. The model uses a hybrid Latent Mixture-of-Experts (LatentMoE) architecture that interleaves Mamba-2, MoE, and attention layers, activating only about 12B parameters per token to deliver significantly higher inference efficiency than dense models of similar scale. NVIDIA positions Nemotron 3 Super as part of its Nemotron 3 family (Nano, Super, Ultra), focusing this tier on multi-agent coordination, long-context reasoning, and high-volume enterprise workloads. Nemotron 3 Super offers a context window of up to 1 million tokens, aimed at long-horizon planning, cross-document reasoning, and multi-user or multi-agent conversations. NVIDIA reports that Super achieves up to ~5x higher throughput than its previous Nemotron models and more than 50% higher token generation speed than leading open peers, helped by multi-token prediction (MTP) and the LatentMoE design, which can consult multiple experts at the cost of one. The model is trained with NVFP4 quantization for compute efficiency and incorporates multi-environment reinforcement learning across more than ten environments to improve reasoning and tool-use performance, with strong results on benchmarks such as AIME 2025, TerminalBench, and SWE-Bench Verified. Strategically, NVIDIA is releasing Nemotron 3 Super as an open model with downloadable weights, data recipes, and a permissive commercial license, under the NVIDIA Nemotron Open Model License, allowing self-hosting, fine-tuning, and on-prem deployment. The model is distributed via platforms such as NVIDIA NIM, major cloud providers, and hosting services including Hugging Face, OpenRouter, and others, and can be run on multi-GPU server configurations (e.g., 8× H100) or, in more optimized/distilled forms, on smaller setups with around 64GB memory for local experimentation. For the broader AI ecosystem, Nemotron 3 Super underscores NVIDIA’s push to pair its GPU and inference stack with high-performance open models optimized for agentic AI, giving enterprises and developers an alternative to closed proprietary models for long-context reasoning, automation, and multi-agent systems.
AI-generated background, compiled from web sources — not editorial content.

𝕏/@JensenHuang ·

𝕏/@claudeai ·

The Block ·

Coindesk ·

𝕏/@sarahdingwang ·

𝕏/@PeterMcCrory ·

𝕏/@JensenHuang ·

𝕏/@claudeai ·

The Block ·

Coindesk ·

𝕏/@sarahdingwang ·

𝕏/@PeterMcCrory ·
🚀 Love DeFi? Ready to dive in and start earning $SQUID while making an impact?