Nvidia Publishes Nemotron 3 Ultra Model Card With H100 and B200 Benchmarks
A model card for Nvidia's Nemotron 3 Ultra is now available, adding benchmarks across TRT-LLM, vLLM and SGLang on H100 and B200 GPUs for the company's 550 billion-parameter open model. It also indicates the model's FP4 weights can run on H100 hardware, giving developers clearer deployment guidance after the initial launch.
Nvidia introduced Nemotron 3 Ultra earlier this week as a mixture-of-experts system for long-running AI agents in coding, tool-use and research workflows. The company said the model has 55 billion active parameters and a 1 million-token context window, and can deliver up to five times faster inference while lowering the cost of complex agentic tasks by as much as 30% versus other open frontier models.
From the sources (25 posts)
@ctnzrNVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premise, on the cloud, or at the edge. Model is live on HuggingFace under the OpenMDW 1.1 license.
@ctnzrAccuracy up there with the best open weight models, but 2-6X faster, thanks to using very little full attention and relying mostly on SSM, as well as LatentMoE. Pretrained in NVFP4. We detail the pretraining process including two divergenc
@charles_irlRT @ctnzr: NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premis…
@thezachmuellerRT @ctnzr: NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premis…
@togethercomputeNemotron 3 Ultra is built for long-horizon agents that need to plan, code, test, debug, and iterate across large working sets. 550B parameters, 55B active, 1M context, Hybrid Mamba-Transformer MoE, Multi-Token Prediction, and multi-environ
@victormustarRT @HuggingPapers: NVIDIA just released Nemotron 3 Ultra on Hugging Face 550B total params, 55B active, hybrid Mamba-2 MoE Transformer, 1…
@charles_irlThe Nemotron series is impressive -- strong capabilities in an efficient form factor, with high-performance implementations (esp on Blackwell) in open source. Deploy the new Nemotron 3 Ultra (550B-A55B-NVFP4) on @modal starting from this r
@nvidiaaiToday we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.
@nvidiaaiUltra excels at complex tasks like coding and deep research. Long-running agents spend their time planning, using tools, recovering from failures, and deciding what to do next. The model’s hybrid Mamba-Transformer MoE architecture enables
@nvidiaaiNemotron 3 Ultra delivers leading accuracy for agentic tasks, including agent productivity, coding, and long horizon planning.
@nvidiaaiBeyond benchmark performance, Ultra can work through large codebases, reason across long chains of tool calls, and synthesize information gathered from hundreds of sources.
@nvidiaaiWe post-trained Ultra for popular agent harnesses like @openclaw, @NousResearch Hermes Agent, and @Langchain. The result is an open frontier model developers can customize for specialized agents across domains. Read more:
@nvidiaai@openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, and post-training recipes. Available now on @huggingface →
@lmsysorg🎉 Meet Nemotron 3 Ultra from @nvidia, a frontier reasoning model, with 550B total params (55B active) built for long-running autonomous agents. Day-0 support is now live in SGLang, plus Day-0 RL support with Miles: GRPO training on 128 H2
@huggingfaceRT @NVIDIAAI: Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delive…
@mervenoyannNVIDIA Nemotron Ultra is here 😍 > 55B/550B a hybrid MoE  🦖 with 1M context window > supports MTP speculative decoding 💨 > day-0 supported in transformers sits in the most attractive quadrant per performance/efficiency in AA Ind
@lmsysorg🚀 New blog: SGLang and Miles Add Day-0 Support for NVIDIA Nemotron 3 Ultra for Long-Running Autonomous Agents The hard part of agentic workloads isn't one big answer, it's sustaining reasoning across hundreds of steps. Here's how we delive
@ctnzrRT @appliedcompute: @nvidia’s Nemotron 3 Ultra handles software-engineering tasks at a fraction of the per-task cost of frontier models. So…
@ctnzrRT @llm_wizard: NEMOTRON 3 ULTRA IS LIVE. OUR BEST MODEL YET. PUNCHING IN THE SAME BALLPARK AS THE OPEN FRONTIER BAYBEEEEE. RECIPES? CHECK…
@artificialanlysNVIDIA has just released Nemotron 3 Ultra, the new most intelligent US open weights model, with leading speed for its intelligence Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, well ahead of the next strongest
@natolambertNvidia joined the multi-teacher, on-policy distillation (MODP) gang! Is industry standard post-training right now. The multi-teacher SFT to RL that Microsoft did in their first model was the standard established by DeepSeek R1. I expect MA
@brianroemmeleNew open source NVIDIA Nemotron 3 Ultra (BF16 weights) is quite good! I am testing it now. It is ranking high and is quite fast. Nvidia is leading the way in the US with powerful open source models. Nemotron v3 models, datasets, etc: h
@basetenAre you tired of waiting 17 minutes for an AI agent to finish a code change? As an agent’s context grows, standard transformer attention can turn long runs into a bottleneck. @NVIDIAAI Nemotron 3 Ultra addresses this with a hybrid archite
@mtsliveSITUATION DETECTED: Nvidia has launched Nemotron 3 Ultra, a 550B open model built for long-running agents. It delivers 5x faster inference and cuts the cost of agentic tasks by up to 30% versus other open frontier models.
@_akhaliqRT @NVIDIAAI: @openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, a…