Command Palette
Search for a command to run...

Nvidia Publishes Nemotron 3 Ultra Model Card With H100 and B200 Benchmarks

aiai-modelingai-model-releasesai-open-modelsai-infrastructureai-inference-platforms 102 posts · 43 accounts

A model card for Nvidia's Nemotron 3 Ultra is now available, adding benchmarks across TRT-LLM, vLLM and SGLang on H100 and B200 GPUs for the company's 550 billion-parameter open model. It also indicates the model's FP4 weights can run on H100 hardware, giving developers clearer deployment guidance after the initial launch.

Nvidia introduced Nemotron 3 Ultra earlier this week as a mixture-of-experts system for long-running AI agents in coding, tool-use and research workflows. The company said the model has 55 billion active parameters and a 1 million-token context window, and can deliver up to five times faster inference while lowering the cost of complex agentic tasks by as much as 30% versus other open frontier models.

From the sources (25 posts)

@ctnzr

NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premise, on the cloud, or at the edge. Model is live on HuggingFace under the OpenMDW 1.1 license.

@ctnzr

Accuracy up there with the best open weight models, but 2-6X faster, thanks to using very little full attention and relying mostly on SSM, as well as LatentMoE. Pretrained in NVFP4. We detail the pretraining process including two divergenc

@charles_irl

RT @ctnzr: NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premis…

@thezachmueller

RT @ctnzr: NVIDIA Nemotron 3 Ultra is now live! Frontier accuracy, 5X greater speed, 30% lower cost. Deploy however you need - on-premis…

@togethercompute

Nemotron 3 Ultra is built for long-horizon agents that need to plan, code, test, debug, and iterate across large working sets. 550B parameters, 55B active, 1M context, Hybrid Mamba-Transformer MoE, Multi-Token Prediction, and multi-environ

@victormustar

RT @HuggingPapers: NVIDIA just released Nemotron 3 Ultra on Hugging Face 550B total params, 55B active, hybrid Mamba-2 MoE Transformer, 1…

@charles_irl

The Nemotron series is impressive -- strong capabilities in an efficient form factor, with high-performance implementations (esp on Blackwell) in open source. Deploy the new Nemotron 3 Ultra (550B-A55B-NVFP4) on @modal starting from this r

@nvidiaai

Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.

@nvidiaai

Ultra excels at complex tasks like coding and deep research. Long-running agents spend their time planning, using tools, recovering from failures, and deciding what to do next. The model’s hybrid Mamba-Transformer MoE architecture enables

@nvidiaai

Nemotron 3 Ultra delivers leading accuracy for agentic tasks, including agent productivity, coding, and long horizon planning.

@nvidiaai

Beyond benchmark performance, Ultra can work through large codebases, reason across long chains of tool calls, and synthesize information gathered from hundreds of sources.

@nvidiaai

We post-trained Ultra for popular agent harnesses like @openclaw, @NousResearch Hermes Agent, and @Langchain. The result is an open frontier model developers can customize for specialized agents across domains. Read more:

@nvidiaai

@openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, and post-training recipes. Available now on @huggingface →

@lmsysorg

🎉 Meet Nemotron 3 Ultra from @nvidia, a frontier reasoning model, with 550B total params (55B active) built for long-running autonomous agents. Day-0 support is now live in SGLang, plus Day-0 RL support with Miles: GRPO training on 128 H2

@huggingface

RT @NVIDIAAI: Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delive…

@mervenoyann

NVIDIA Nemotron Ultra is here 😍 > 55B/550B a hybrid MoE  🦖 with 1M context window > supports MTP speculative decoding 💨 > day-0 supported in transformers sits in the most attractive quadrant per performance/efficiency in AA Ind

@lmsysorg

🚀 New blog: SGLang and Miles Add Day-0 Support for NVIDIA Nemotron 3 Ultra for Long-Running Autonomous Agents The hard part of agentic workloads isn't one big answer, it's sustaining reasoning across hundreds of steps. Here's how we delive

@ctnzr

RT @appliedcompute: @nvidia’s Nemotron 3 Ultra handles software-engineering tasks at a fraction of the per-task cost of frontier models. So…

@ctnzr

RT @llm_wizard: NEMOTRON 3 ULTRA IS LIVE. OUR BEST MODEL YET. PUNCHING IN THE SAME BALLPARK AS THE OPEN FRONTIER BAYBEEEEE. RECIPES? CHECK…

@artificialanlys

NVIDIA has just released Nemotron 3 Ultra, the new most intelligent US open weights model, with leading speed for its intelligence Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, well ahead of the next strongest

@natolambert

Nvidia joined the multi-teacher, on-policy distillation (MODP) gang! Is industry standard post-training right now. The multi-teacher SFT to RL that Microsoft did in their first model was the standard established by DeepSeek R1. I expect MA

@brianroemmele

New open source NVIDIA Nemotron 3 Ultra (BF16 weights) is quite good! I am testing it now. It is ranking high and is quite fast. Nvidia is leading the way in the US with powerful open source models. Nemotron v3 models, datasets, etc: h

@baseten

Are you tired of waiting 17 minutes for an AI agent to finish a code change? As an agent’s context grows, standard transformer attention can turn long runs into a bottleneck. @NVIDIAAI Nemotron 3 Ultra addresses this with a hybrid archite

@mtslive

SITUATION DETECTED: Nvidia has launched Nemotron 3 Ultra, a 550B open model built for long-running agents. It delivers 5x faster inference and cuts the cost of agentic tasks by up to 30% versus other open frontier models.

@_akhaliq

RT @NVIDIAAI: @openclaw @NousResearch @LangChain As always, Nemotron 3 Ultra is fully open. This includes model weights, synthetic data, a…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive