Command Palette
Search for a command to run...

Nvidia's 4-Step Distilled Models Top Artificial Analysis Open-Weight Image to Video Leaderboard

aiai-modelingai-model-releasesai-open-modelsai-research-evals 2 posts · 1 accounts

Nvidia's distilled variants of its Cosmos3-Super model have reduced inference steps from 35 to just 4 for image to video generation, removing the need for classifier-free guidance. The Cosmos3-Super-Image2Video-4Step model now ranks as the top open-weights model in Artificial Analysis's image to video leaderboard, surpassing the full 64B version.

Nvidia positioned the Cosmos 3 family as foundation models for physical AI, targeting training data and video generation for robotics, autonomous vehicles, and industrial systems. The models, originally released with open weights on July 20, produce up to 720p video clips averaging 8 seconds long and are available on Hugging Face under a commercial-use license. The text to image variant of the model ranks third on the same platform's leaderboard.

From the sources (2 posts)

@artificialanlys

NVIDIA's Cosmos3-Super-4Step distills are the new #1 open weights model for Image to Video and #3 for Text to Image in the Artificial Analysis Arena Cosmos3-Super-Text2Image-4Step and Cosmos3-Super-Image2Video-4Step are distilled variants

@artificialanlys

Check out the Cosmos3-Super-4Step models for yourself on the Artificial Analysis Leaderboards: Text to Image: Image to Video: Or vote in the Image Arena: And the Vi

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive