Z.ai's Open-Weight GLM-5.2 Ties Claude Opus 4.8 at 20.9% on CritPt
Z.ai's GLM-5.2 is strengthening its position as a leading open-weight AI model after new benchmark results showed its max-reasoning configuration matching Anthropic's Claude Opus 4.8 on CritPt, a test of unpublished research-level physics problems, at 20.9%. Artificial Analysis said the score put GLM-5.2 well ahead of the next-best open model, DeepSeek V4 Pro at 12.9%, while Vals ranked it the top open-weight system on the Vals Index, Vibe Code Bench and Terminal Bench. It also reached first place on Design Arena with a 1,360 Elo, up 27 points.
The MIT-licensed model, which has a 1 million-token context window and the same pricing as GLM-5.1, also scored 51 on Artificial Analysis's Intelligence Index and 1,524 on GDPval-AA v2, underscoring its competitiveness on longer agentic workloads. Artificial Analysis said CritPt is independently benchmarked using privately graded problems developed by Argonne and UIUC with contributions from more than 60 researchers, and even GPT-5.5 Pro solves under a third of the questions. GLM-5.2 is also spreading across deployment options, including Go support and private encrypted access on Venice.
From the sources (25 posts)
@zai_orgIntroducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits,
@zai_orgFor GLM-5.2, we strengthened 1M-context training for coding agents across large-scale implementation, automated research, performance optimization, and complex debugging. The result is a long-context system that is both broad in scope and r
@zephyr_z9Zhipu is about to drop a banger
@zephyr_z9HOLY SHIT!!
@victormustarGLM-5.2 is available on Hugging Face 🔥 It's an important day for open source AI: opus-class frontier intelligence, 1M context, agentic-first by design. --> The future of AI and humanity is open
@clementdelangueRT @Zai_org: Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long…
@zephyr_z9Zhipu did an absolutely crazy drop
@teortaxestex> We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. We also improve GLM-5.2’s MTP Good demonstration that a lot could be squeezed o
@testingcatalogZAI 🔥: GLM-5.2 is now available on huggingface! > It comes with a 1M context window and 2 levels of reasoning effort, max and high. MIT license and same pricing as GLM-5.1. > GLM-5.2 scores 46.2% on DeepSWE, the SOTA score among op
@lmsysorg🎉 Meet GLM-5.2 from @Zai_org, the new flagship for long-horizon tasks built on a solid 1M-token context. Day-0 support is now live in SGLang! ✅ Advanced coding capability: 81.0 on Terminal-Bench 2.1 (vs 62.0 for GLM-5.1) ✅ Solid 1M contex
@calebfahlgrenRT @Zai_org: Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long…
@zai_orgRT @ProximalHQ: GLM 5.2 ranks #3 on FrontierSWE. It is only behind Fable 5 and Opus 4.8, and it outperforms GPT-5.5. This is the first mod…
@zai_orgRT @Designarena: BREAKING: GLM-5.2 is now 1st on Design Arena. With an Elo of 1360, GLM-5.2 has jumped ahead of the now unavailable Claude…
@teortaxestexGLM 5.2 is apparently the strongest Chinese model, and it'll be open this is quite a jump on CritPt. They have *finally* started to build general reasoners. I think it'll be very competitive on ARC-AGI 2 and WeirdML, too
@zephyr_z9RT @teortaxesTex: GLM 5.2 is apparently the strongest Chinese model, and it'll be open this is quite a jump on CritPt. They have *finally*…
@zephyr_z9Zhipu has the best post-training stack in China
@arenaGLM-5.2 (Max) by @Zai_org ranks #10 on the new Agent Arena leaderboard, closely matching Claude-Opus-4.8 (non-thinking) and is the #1 open model by a wide margin! In Agent Arena, we measure models on millions of real-world, long-horizon ag
@victormustarhuge
@adinayakupGLM 5.2 is here 🔥 ✨ 753B ( smaller than you expect? 👀) ✨ 1M context ✨ MIT license ✨ GLM IndexShare: reuses the indexer across layers, 2.9x fewer FLOPs/token at 1M ✨ AIME 2026: 99.2 (beats GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.8 ✨ vLL
@sentdexZai was gracious enough to give me a key to test out GLM 5.2. I used it on a few simple tasks and quickly realized this model is on another level. I committed to using GLM 5.2 solely for the weekend and yesterday on everything from simple
@thezachmuellerRT @Zai_org: Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long…
@zephyr_z9@ChrissGPT Here
@kimmonismusLets go, GLM-5.2 released as Open Weights model. tl;dr -1M context window -MIT-licensed open weights -Stronger long-horizon coding agents -Two reasoning modes: max and high -Same API pricing as GLM-5.1 Zai says GLM-5.2 was trained specif
@andrewcurran_Extremely powerful and MIT open-source. Mutuals who had early access are giving it great reviews.
@nrehiew_It looks like GLM 5.2 is the second best long horizon model available??