Command Palette
Search for a command to run...

Nvidia Releases 550B Nemotron 3 Ultra Open Model With 5x Faster Inference

aiai-modelingai-model-releasesai-open-models 101 posts · 43 accounts

Nvidia released Nemotron 3 Ultra, a 550 billion-parameter mixture-of-experts open model for long-running AI agents used in coding, tool use and research workflows. The company said the model delivers up to five times faster inference and can lower the cost of complex agentic tasks by as much as 30% versus other open frontier models, using a hybrid Mamba-Transformer architecture designed to keep long-context workloads efficient.

Nemotron 3 Ultra has 55 billion active parameters, a 1 million-token context window and was pretrained in NVFP4 on 20 trillion tokens. The release includes base, post-trained and reward checkpoints, along with training data and recipes, while vLLM said it had day-one support. Artificial Analysis, which said it worked with Nvidia ahead of the launch, scored the model at 47.7 on its Intelligence Index, ahead of Gemma 4 31B, Nemotron 3 Super and gpt-oss-120b among U.S. open-weight peers.

From the sources (25 posts)

@ctnzr

NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premise, on the cloud, or at the edge. Model is live on HuggingFace under the OpenMDW 1.1 license.

@ctnzr

Accuracy up there with the best open weight models, but 2-6X faster, thanks to using very little full attention and relying mostly on SSM, as well as LatentMoE. Pretrained in NVFP4. We detail the pretraining process including two divergenc

@charles_irl

RT @ctnzr: NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premis…

@thezachmueller

RT @ctnzr: NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premis…

@togethercompute

Nemotron 3 Ultra is built for long-horizon agents that need to plan, code, test, debug, and iterate across large working sets. 550B parameters, 55B active, 1M context, Hybrid Mamba-Transformer MoE, Multi-Token Prediction, and multi-environ

@victormustar

RT @HuggingPapers: NVIDIA just released Nemotron 3 Ultra on Hugging Face 550B total params, 55B active, hybrid Mamba-2 MoE Transformer, 1…

@charles_irl

The Nemotron series is impressive -- strong capabilities in an efficient form factor, with high-performance implementations (esp on Blackwell) in open source. Deploy the new Nemotron 3 Ultra (550B-A55B-NVFP4) on @modal starting from this r

@nvidiaai

Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.

@nvidiaai

Ultra excels at complex tasks like coding and deep research. Long-running agents spend their time planning, using tools, recovering from failures, and deciding what to do next. The model’s hybrid Mamba-Transformer MoE architecture enables

@nvidiaai

Nemotron 3 Ultra delivers leading accuracy for agentic tasks, including agent productivity, coding, and long horizon planning.

@nvidiaai

Beyond benchmark performance, Ultra can work through large codebases, reason across long chains of tool calls, and synthesize information gathered from hundreds of sources.

@nvidiaai

We post-trained Ultra for popular agent harnesses like @openclaw, @NousResearch Hermes Agent, and @Langchain. The result is an open frontier model developers can customize for specialized agents across domains. Read more:

@nvidiaai

@openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, and post-training recipes. Available now on @huggingface →

@lmsysorg

🎉 Meet Nemotron 3 Ultra from @nvidia, a frontier reasoning model, with 550B total params (55B active) built for long-running autonomous agents. Day-0 support is now live in SGLang, plus Day-0 RL support with Miles: GRPO training on 128 H2

@huggingface

RT @NVIDIAAI: Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delive…

@mervenoyann

NVIDIA Nemotron Ultra is here 😍 > 55B/550B a hybrid MoE  🦖 with 1M context window > supports MTP speculative decoding 💨 > day-0 supported in transformers sits in the most attractive quadrant per performance/efficiency in AA Ind

@lmsysorg

🚀 New blog: SGLang and Miles Add Day-0 Support for NVIDIA Nemotron 3 Ultra for Long-Running Autonomous Agents The hard part of agentic workloads isn't one big answer, it's sustaining reasoning across hundreds of steps. Here's how we delive

@ctnzr

RT @appliedcompute: @nvidia’s Nemotron 3 Ultra handles software-engineering tasks at a fraction of the per-task cost of frontier models. So…

@ctnzr

RT @llm_wizard: NEMOTRON 3 ULTRA IS LIVE. OUR BEST MODEL YET. PUNCHING IN THE SAME BALLPARK AS THE OPEN FRONTIER BAYBEEEEE. RECIPES? CHECK…

@artificialanlys

NVIDIA has just released Nemotron 3 Ultra, the new most intelligent US open weights model, with leading speed for its intelligence Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, well ahead of the next strongest

@natolambert

Nvidia joined the multi-teacher, on-policy distillation (MODP) gang! Is industry standard post-training right now. The multi-teacher SFT to RL that Microsoft did in their first model was the standard established by DeepSeek R1. I expect MA

@brianroemmele

New open source NVIDIA Nemotron 3 Ultra (BF16 weights) is quite good! I am testing it now. It is ranking high and is quite fast. Nvidia is leading the way in the US with powerful open source models. Nemotron v3 models, datasets, etc: h

@baseten

Are you tired of waiting 17 minutes for an AI agent to finish a code change? As an agent’s context grows, standard transformer attention can turn long runs into a bottleneck. @NVIDIAAI Nemotron 3 Ultra addresses this with a hybrid archite

@mtslive

SITUATION DETECTED: Nvidia has launched Nemotron 3 Ultra, a 550B open model built for long-running agents. It delivers 5x faster inference and cuts the cost of agentic tasks by up to 30% versus other open frontier models.

@_akhaliq

RT @NVIDIAAI: @openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, a…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive