Command Palette
Search for a command to run...

NVIDIA Releases NVFP4 Checkpoint for GLM 5.2 AI Model

aiai-modelingai-model-releasesai-infrastructureai-inference-platformsai-open-models 21 posts · 19 accounts

NVIDIA released an official NVFP4 checkpoint for GLM 5.2, a large language model developed by Z.ai. The quantized checkpoint optimizes the model for NVIDIA's Blackwell GPUs, reducing the memory footprint compared to standard FP8 precision while maintaining accuracy across coding and reasoning benchmarks.

NVFP4 is an efficient tensor format designed to run frontier AI models on hardware with reduced memory requirements. The checkpoint is now available for download and supports immediate inference through popular open-source frameworks like vLLM and SGLang. The release follows strong adoption of GLM 5.2 since its June debut, where it quickly ranked among the top performers in coding competitions and AI coding assistants.

From the sources (21 posts)

@fchollet

This is the strongest ARC-AGI-2 performance to date by an open-source model.

@rauchg

Really fast GLM now live

@arena

How did GLM-5.2 (Max) get to the top of Code Arena: Frontend? Looking at matched head-to-head on real-world web dev frontend tasks, @Zai_org's latest model takes a higher win share than its opponent in every pairing but one. - Beats every

@leerob

You can now try GLM 5.2 in Cursor! Excited to see more useful open models, thank you to Fireworks for partnering here. Results from our evals ↓

@eliebakouch

wow GLM5.2 is at the same level as opus 4.8 in terms of cost efficiency on cursorbench

@scaling01

Sonnet 5 can't come soon enough

@cognition

Try Kimi K2.7 and GLM 5.2 for free in Devin Desktop and CLI

@zai_org

RT @ZixuanLi_: GLM-5.2 is now available in Cursor. The model has performed strongly on OpenRouter's Cursor usage rankings over the past we…

@dabit3

Two of the strongest new coding models are free to try in Devin. Kimi K2.7 Code and GLM 5.2 in Devin CLI and Devin Desktop, free for Pro/Max/Teams users.

@imjaredz

RT @cognition: Try Kimi K2.7 and GLM 5.2 for free in Devin Desktop and CLI

@zai_org

RT @ZixuanLi_: Right alongside Cursor, Devin Desktop (Windsurf) and CLI now support GLM-5.2 as well. FrontierCode Extended is a benchmark…

@thezachmueller

RT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…

@_akhaliq

RT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…

@huggingpapers

NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for Blackwell GPUs— nearly matching FP8 accuracy.

@teortaxestex

RT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…

@lmsysorg

🎉 NVIDIA just released an NVFP4 checkpoint of GLM-5.2 from @Zai_org, a 744B MoE (40B active) for reasoning & coding. Day-0 support is live in SGLang! 🤝 @nvidia > NVFP4 quantization via NVIDIA Model Optimizer: frontier-class reasoning at a

@_akhaliq

RT @ZixuanLi_: The official GLM-5.2 NVFP4 from NVIDIA is now available. Curious how it compares to other quantizations.

@clementdelangue

RT @ZixuanLi_: The official GLM-5.2 NVFP4 from NVIDIA is now available. Curious how it compares to other quantizations.

@charles_irl

RT @lmsysorg: 🎉 NVIDIA just released an NVFP4 checkpoint of GLM-5.2 from @Zai_org, a 744B MoE (40B active) for reasoning & coding. Day-0 su…

@vllm_project

GLM-5.2 in NVFP4 is ready to serve in vLLM 🚀 @NVIDIAAI's official NVFP4 checkpoint of GLM-5.2 on Blackwell cuts the memory footprint vs FP8 while matching its accuracy across reasoning, coding, and long-context benchmarks. Serve it today

@zixuanli_

The official GLM-5.2 NVFP4 from NVIDIA is now available. Curious how it compares to other quantizations.

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive