Nvidia Releases 550B Nemotron 3 Ultra Open Model With 5x Faster Inference
Nvidia released Nemotron 3 Ultra, a 550 billion-parameter mixture-of-experts open model for long-running AI agents used in coding, tool use and research workflows. The company said the model delivers up to five times faster inference and can lower the cost of complex agentic tasks by as much as 30% versus other open frontier models, using a hybrid Mamba-Transformer architecture designed to keep long-context workloads efficient.
Nemotron 3 Ultra has 55 billion active parameters, a 1 million-token context window and was pretrained in NVFP4 on 20 trillion tokens. The release includes base, post-trained and reward checkpoints, along with training data and recipes, while vLLM said it had day-one support. Artificial Analysis, which said it worked with Nvidia ahead of the launch, scored the model at 47.7 on its Intelligence Index, ahead of Gemma 4 31B, Nemotron 3 Super and gpt-oss-120b among U.S. open-weight peers.
From the sources (25 posts)
@ctnzrNVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premise, on the cloud, or at the edge. Model is live on HuggingFace under the OpenMDW 1.1 license.
@ctnzrAccuracy up there with the best open weight models, but 2-6X faster, thanks to using very little full attention and relying mostly on SSM, as well as LatentMoE. Pretrained in NVFP4. We detail the pretraining process including two divergenc
@charles_irlRT @ctnzr: NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premis…
@thezachmuellerRT @ctnzr: NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premis…
@togethercomputeNemotron 3 Ultra is built for long-horizon agents that need to plan, code, test, debug, and iterate across large working sets. 550B parameters, 55B active, 1M context, Hybrid Mamba-Transformer MoE, Multi-Token Prediction, and multi-environ
@victormustarRT @HuggingPapers: NVIDIA just released Nemotron 3 Ultra on Hugging Face 550B total params, 55B active, hybrid Mamba-2 MoE Transformer, 1…
@charles_irlThe Nemotron series is impressive -- strong capabilities in an efficient form factor, with high-performance implementations (esp on Blackwell) in open source. Deploy the new Nemotron 3 Ultra (550B-A55B-NVFP4) on @modal starting from this r
@nvidiaaiToday we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.
@nvidiaaiUltra excels at complex tasks like coding and deep research. Long-running agents spend their time planning, using tools, recovering from failures, and deciding what to do next. The model’s hybrid Mamba-Transformer MoE architecture enables
@nvidiaaiNemotron 3 Ultra delivers leading accuracy for agentic tasks, including agent productivity, coding, and long horizon planning.
@nvidiaaiBeyond benchmark performance, Ultra can work through large codebases, reason across long chains of tool calls, and synthesize information gathered from hundreds of sources.
@nvidiaaiWe post-trained Ultra for popular agent harnesses like @openclaw, @NousResearch Hermes Agent, and @Langchain. The result is an open frontier model developers can customize for specialized agents across domains. Read more:
@nvidiaai@openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, and post-training recipes. Available now on @huggingface →
@lmsysorg🎉 Meet Nemotron 3 Ultra from @nvidia, a frontier reasoning model, with 550B total params (55B active) built for long-running autonomous agents. Day-0 support is now live in SGLang, plus Day-0 RL support with Miles: GRPO training on 128 H2
@huggingfaceRT @NVIDIAAI: Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delive…
@mervenoyannNVIDIA Nemotron Ultra is here 😍 > 55B/550B a hybrid MoE  🦖 with 1M context window > supports MTP speculative decoding 💨 > day-0 supported in transformers sits in the most attractive quadrant per performance/efficiency in AA Ind
@lmsysorg🚀 New blog: SGLang and Miles Add Day-0 Support for NVIDIA Nemotron 3 Ultra for Long-Running Autonomous Agents The hard part of agentic workloads isn't one big answer, it's sustaining reasoning across hundreds of steps. Here's how we delive
@ctnzrRT @appliedcompute: @nvidia’s Nemotron 3 Ultra handles software-engineering tasks at a fraction of the per-task cost of frontier models. So…
@ctnzrRT @llm_wizard: NEMOTRON 3 ULTRA IS LIVE. OUR BEST MODEL YET. PUNCHING IN THE SAME BALLPARK AS THE OPEN FRONTIER BAYBEEEEE. RECIPES? CHECK…
@artificialanlysNVIDIA has just released Nemotron 3 Ultra, the new most intelligent US open weights model, with leading speed for its intelligence Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, well ahead of the next strongest
@natolambertNvidia joined the multi-teacher, on-policy distillation (MODP) gang! Is industry standard post-training right now. The multi-teacher SFT to RL that Microsoft did in their first model was the standard established by DeepSeek R1. I expect MA
@brianroemmeleNew open source NVIDIA Nemotron 3 Ultra (BF16 weights) is quite good! I am testing it now. It is ranking high and is quite fast. Nvidia is leading the way in the US with powerful open source models. Nemotron v3 models, datasets, etc: h
@basetenAre you tired of waiting 17 minutes for an AI agent to finish a code change? As an agent’s context grows, standard transformer attention can turn long runs into a bottleneck. @NVIDIAAI Nemotron 3 Ultra addresses this with a hybrid archite
@mtsliveSITUATION DETECTED: Nvidia has launched Nemotron 3 Ultra, a 550B open model built for long-running agents. It delivers 5x faster inference and cuts the cost of agentic tasks by up to 30% versus other open frontier models.
@_akhaliqRT @NVIDIAAI: @openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, a…