Nvidia Releases 3B-to-14B Nemotron-Labs-Diffusion Language Models for Faster GPU Inference
Nvidia released Nemotron-Labs-Diffusion, a family of diffusion language models that it said generates multiple tokens in parallel within a single model rather than one token at a time. The company said the models can revise tokens as they go instead of committing to each token permanently, a design it said delivers faster inference and better use of modern GPUs.
The model family ranges from 3B to 14B and includes vision-language variants, which Nvidia said are available now.
From the sources (4 posts)
@emostaqueRT @PavloMolchanov: We’re releasing Nemotron-Labs-Diffusion - the first Tri-mode LM family (3B/8B/14B) that switches between 1⃣Autoregressi…
@nvidiaaiMost language models only generate one token at a time. We just released Nemotron-Labs-Diffusion, a family of diffusion language models that take a different approach, generating multiple tokens in parallel within a single model. Rather th
@huggingfaceRT @NVIDIAAI: Most language models only generate one token at a time. We just released Nemotron-Labs-Diffusion, a family of diffusion lang…
@andrew_n_carrRT @NVIDIAAI: Most language models only generate one token at a time. We just released Nemotron-Labs-Diffusion, a family of diffusion lang…