NVIDIA Releases NVFP4 Checkpoint for GLM 5.2 AI Model
NVIDIA released an official NVFP4 checkpoint for GLM 5.2, a large language model developed by Z.ai. The quantized checkpoint optimizes the model for NVIDIA's Blackwell GPUs, reducing the memory footprint compared to standard FP8 precision while maintaining accuracy across coding and reasoning benchmarks.
NVFP4 is an efficient tensor format designed to run frontier AI models on hardware with reduced memory requirements. The checkpoint is now available for download and supports immediate inference through popular open-source frameworks like vLLM and SGLang. The release follows strong adoption of GLM 5.2 since its June debut, where it quickly ranked among the top performers in coding competitions and AI coding assistants.
From the sources (21 posts)
@fcholletThis is the strongest ARC-AGI-2 performance to date by an open-source model.
@rauchgReally fast GLM now live
@arenaHow did GLM-5.2 (Max) get to the top of Code Arena: Frontend? Looking at matched head-to-head on real-world web dev frontend tasks, @Zai_org's latest model takes a higher win share than its opponent in every pairing but one. - Beats every
@leerobYou can now try GLM 5.2 in Cursor! Excited to see more useful open models, thank you to Fireworks for partnering here. Results from our evals ↓
@eliebakouchwow GLM5.2 is at the same level as opus 4.8 in terms of cost efficiency on cursorbench
@scaling01Sonnet 5 can't come soon enough
@cognitionTry Kimi K2.7 and GLM 5.2 for free in Devin Desktop and CLI
@zai_orgRT @ZixuanLi_: GLM-5.2 is now available in Cursor. The model has performed strongly on OpenRouter's Cursor usage rankings over the past we…
@dabit3Two of the strongest new coding models are free to try in Devin. Kimi K2.7 Code and GLM 5.2 in Devin CLI and Devin Desktop, free for Pro/Max/Teams users.
@imjaredzRT @cognition: Try Kimi K2.7 and GLM 5.2 for free in Devin Desktop and CLI
@zai_orgRT @ZixuanLi_: Right alongside Cursor, Devin Desktop (Windsurf) and CLI now support GLM-5.2 as well. FrontierCode Extended is a benchmark…
@thezachmuellerRT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…
@_akhaliqRT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…
@huggingpapersNVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for Blackwell GPUs— nearly matching FP8 accuracy.
@teortaxestexRT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…
@lmsysorg🎉 NVIDIA just released an NVFP4 checkpoint of GLM-5.2 from @Zai_org, a 744B MoE (40B active) for reasoning & coding. Day-0 support is live in SGLang! 🤝 @nvidia > NVFP4 quantization via NVIDIA Model Optimizer: frontier-class reasoning at a
@_akhaliqRT @ZixuanLi_: The official GLM-5.2 NVFP4 from NVIDIA is now available. Curious how it compares to other quantizations.
@clementdelangueRT @ZixuanLi_: The official GLM-5.2 NVFP4 from NVIDIA is now available. Curious how it compares to other quantizations.
@charles_irlRT @lmsysorg: 🎉 NVIDIA just released an NVFP4 checkpoint of GLM-5.2 from @Zai_org, a 744B MoE (40B active) for reasoning & coding. Day-0 su…
@vllm_projectGLM-5.2 in NVFP4 is ready to serve in vLLM 🚀 @NVIDIAAI's official NVFP4 checkpoint of GLM-5.2 on Blackwell cuts the memory footprint vs FP8 while matching its accuracy across reasoning, coding, and long-context benchmarks. Serve it today
@zixuanli_The official GLM-5.2 NVFP4 from NVIDIA is now available. Curious how it compares to other quantizations.