Command Palette
Search for a command to run...

Nvidia Releases 3B-to-14B Nemotron-Labs-Diffusion Language Models for Faster GPU Inference

aiai-modelingai-model-releases 4 posts · 4 accounts

Nvidia released Nemotron-Labs-Diffusion, a family of diffusion language models that it said generates multiple tokens in parallel within a single model rather than one token at a time. The company said the models can revise tokens as they go instead of committing to each token permanently, a design it said delivers faster inference and better use of modern GPUs.

The model family ranges from 3B to 14B and includes vision-language variants, which Nvidia said are available now.

From the sources (4 posts)

@emostaque

RT @PavloMolchanov: We’re releasing Nemotron-Labs-Diffusion - the first Tri-mode LM family (3B/8B/14B) that switches between 1⃣Autoregressi…

@nvidiaai

Most language models only generate one token at a time. We just released Nemotron-Labs-Diffusion, a family of diffusion language models that take a different approach, generating multiple tokens in parallel within a single model. Rather th

@huggingface

RT @NVIDIAAI: Most language models only generate one token at a time. We just released Nemotron-Labs-Diffusion, a family of diffusion lang…

@andrew_n_carr

RT @NVIDIAAI: Most language models only generate one token at a time. We just released Nemotron-Labs-Diffusion, a family of diffusion lang…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive