NVIDIA Publishes Optimized GLM 5.2 Model on Hugging Face
NVIDIA has published an optimized version of GLM-5.2 on Hugging Face, the open repository for artificial intelligence models. The architecture is a 753 billion parameter mixture of experts (MoE) model developed by Beijing-based AI firm Z.ai, featuring a 1 million token context window. The released package supports NVFP4 quantization, a compression format designed for NVIDIA's Blackwell GPU architecture that maintains performance near standard FP8 precision.
The publication on Hugging Face makes the model directly accessible to developers without relying on intermediary hosting or inference providers. The release expands GLM-5.2's distribution following its broader integration into software platforms earlier this week, including the Devin coding agent, Cursor IDE, and Perplexity's search API. According to community benchmark rankings tracked by The Arena, GLM-5.2 leads the Frontend Coding Arena, achieving higher win rates than rival models from Anthropic and OpenAI in real-world web development tasks.
From the sources (17 posts)
@denisyaratsRT @perplexitydevs: @Zai_org's flagship model, GLM-5.2, is now available in Perplexity's Agent API. GLM-5.2 is one of the strongest open-s…
@zai_orgRT @ZixuanLi_: GLM-5.2 is available in Perplexity's Agent API. Just tested it, and it's powerful when paired with the Search SDK inside a s…
@aravsrinivasGLM now supported on Perplexity Agent API
@fcholletThis is the strongest ARC-AGI-2 performance to date by an open-source model.
@rauchgReally fast GLM now live
@arenaHow did GLM-5.2 (Max) get to the top of Code Arena: Frontend? Looking at matched head-to-head on real-world web dev frontend tasks, @Zai_org's latest model takes a higher win share than its opponent in every pairing but one. - Beats every
@leerobYou can now try GLM 5.2 in Cursor! Excited to see more useful open models, thank you to Fireworks for partnering here. Results from our evals ↓
@eliebakouchwow GLM5.2 is at the same level as opus 4.8 in terms of cost efficiency on cursorbench
@scaling01Sonnet 5 can't come soon enough
@cognitionTry Kimi K2.7 and GLM 5.2 for free in Devin Desktop and CLI
@zai_orgRT @ZixuanLi_: GLM-5.2 is now available in Cursor. The model has performed strongly on OpenRouter's Cursor usage rankings over the past we…
@dabit3Two of the strongest new coding models are free to try in Devin. Kimi K2.7 Code and GLM 5.2 in Devin CLI and Devin Desktop, free for Pro/Max/Teams users.
@imjaredzRT @cognition: Try Kimi K2.7 and GLM 5.2 for free in Devin Desktop and CLI
@zai_orgRT @ZixuanLi_: Right alongside Cursor, Devin Desktop (Windsurf) and CLI now support GLM-5.2 as well. FrontierCode Extended is a benchmark…
@thezachmuellerRT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…
@_akhaliqRT @HuggingPapers: NVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for…
@huggingpapersNVIDIA just released an optimized GLM-5.2 on Hugging Face A 753B parameter MoE with 1M context, quantized to NVFP4 for Blackwell GPUs— nearly matching FP8 accuracy.