Nvidia Releases 550B Open-Weight Nemotron 3 Ultra With 1M-Token Context
Nvidia released Nemotron 3 Ultra, a 550 billion-parameter open-weight model built for long-running autonomous agents, with 55 billion active parameters per forward pass, a 1 million-token context window and up to 5.9 times the throughput of competing open mixture-of-experts models.
The release adds detail to earlier results from Harvey and Trajectory Labs, which said they post-trained Nemotron 3 Ultra for legal work less than 24 hours after launch. Harvey said the post-trained system’s all-pass rate on Harvey’s Legal Agent Benchmark rose to 5.8% from 0%, placing it between Sonnet 4.6 at 4.2% and Opus 4.6 at 6.6%, while lowering per-token costs to roughly one-eighth to one-fiftieth of those closed models.
From the sources (8 posts)
@nvidiaaiRT @harvey: We partnered with @trajectorylabs to post-train NVIDIA Nemotron 3 Ultra for legal. Here’s what we found: 1) Open-weight models…
@ctnzrRT @harvey: We partnered with @trajectorylabs to post-train NVIDIA Nemotron 3 Ultra for legal. Here’s what we found: 1) Open-weight models…
@harveyWe partnered with @trajectorylabs to post-train NVIDIA Nemotron 3 Ultra for legal. Here’s what we found: 1) Open-weight models can reach frontier legal performance. On our Legal Agent Benchmark (LAB), Nemotron 3 Ultra started at a 0% all-
@trajectorylabs1/ We post-trained @nvidia Nemotron 3 Ultra on @harvey Legal Agent Bench in under 24 hours. The result: an open model reaching the same band as leading closed models on legal work, at a fraction of the cost. The correlating story: when a
@dl_weekly🤖 From this week's issue: NVIDIA releases Nemotron 3 Ultra — 550B total / 55B active MoE hybrid Mamba-Transformer, pretrained in NVFP4, with up to 5.9x throughput over competing open MoEs and 1M token context.
@tekniumRT @leopardracer: NVIDIA’S FREE 550B MODEL JUST DROPPED FOR HERMES AGENT nemotron 3 ultra is 5x faster and built for long-running agents,…
@idleprotocol(2/2) What Nemotron 3 Ultra actually is. 550 billion parameters. 55 billion active per forward pass via mixture-of-experts routing. 300+ tokens per second. 1 million token context window. Intelligence Index score of 48 - the highest ever
@leopardracerNVIDIA’S FREE 550B MODEL JUST DROPPED FOR HERMES AGENT nemotron 3 ultra is 5x faster and built for long-running agents, free for 2 weeks on nous portal this guy handed hermes his entire business for 3 weeks and woke up to a completed comp