NVIDIA Releases NVFP4 Checkpoint of 744B Parameter GLM-5.2 Model
NVIDIA released an optimized checkpoint of GLM-5.2, a 744-billion-parameter mixture-of-experts model from Beijing-based AI lab Z.ai, designed for NVIDIA’s Blackwell and Grace Blackwell GPUs. The update introduces NVFP4 quantization via the NVIDIA Model Optimizer, significantly reducing memory overhead while preserving accuracy close to the model’s full FP8 precision. The quantized checkpoint ships with day-zero support in SGLang, an open-source model serving engine, and incorporates sparse attention mechanisms to efficiently handle context windows up to 1 million tokens.
The hardware optimization follows a wave of toolchain integrations and benchmark validation for GLM-5.2 since its launch earlier this month. The model currently holds the strongest performance record among open-weight alternatives on the ARC-AGI-2 general reasoning benchmark and matches the inference cost efficiency of Anthropic’s Opus 4.8 on developer coding workloads. NVIDIA posted the NVFP4 variant on Hugging Face for immediate download, expanding access to frontier-class reasoning and coding capabilities for enterprise and developer environments.
From the sources (16 posts)
@fcholletThis is the strongest ARC-AGI-2 performance to date by an open-source model.
@rauchgReally fast GLM now live
@arenaHow did GLM-5.2 (Max) get to the top of Code Arena: Frontend? Looking at matched head-to-head on real-world web dev frontend tasks, @Zai_org's latest model takes a higher win share than its opponent in every pairing but one. - Beats every
@leerobYou can now try GLM 5.2 in Cursor! Excited to see more useful open models, thank you to Fireworks for partnering here. Results from our evals ↓
@eliebakouchwow GLM5.2 is at the same level as opus 4.8 in terms of cost efficiency on cursorbench
@scaling01Sonnet 5 can't come soon enough
@cognitionTry Kimi K2.7 and GLM 5.2 for free in Devin Desktop and CLI
@zai_orgRT @ZixuanLi_: GLM-5.2 is now available in Cursor. The model has performed strongly on OpenRouter's Cursor usage rankings over the past we…
@dabit3Two of the strongest new coding models are free to try in Devin. Kimi K2.7 Code and GLM 5.2 in Devin CLI and Devin Desktop, free for Pro/Max/Teams users.
@imjaredzRT @cognition: Try Kimi K2.7 and GLM 5.2 for free in Devin Desktop and CLI
@zai_orgRT @ZixuanLi_: Right alongside Cursor, Devin Desktop (Windsurf) and CLI now support GLM-5.2 as well. FrontierCode Extended is a benchmark…
@thezachmuellerRT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…
@_akhaliqRT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…
@huggingpapersNVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for Blackwell GPUs— nearly matching FP8 accuracy.
@teortaxestexRT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…
@lmsysorg🎉 NVIDIA just released an NVFP4 checkpoint of GLM-5.2 from @Zai_org, a 744B MoE (40B active) for reasoning & coding. Day-0 support is live in SGLang! 🤝 @nvidia > NVFP4 quantization via NVIDIA Model Optimizer: frontier-class reasoning at a