Command Palette
Search for a command to run...

Z.ai's Open-Weight GLM-5.2 Ties Claude Opus 4.8 at 20.9% on CritPt

aiai-modelingai-model-releasesai-open-modelsai-research-evals 105 posts · 49 accounts

Z.ai's GLM-5.2 is strengthening its position as a leading open-weight AI model after new benchmark results showed its max-reasoning configuration matching Anthropic's Claude Opus 4.8 on CritPt, a test of unpublished research-level physics problems, at 20.9%. Artificial Analysis said the score put GLM-5.2 well ahead of the next-best open model, DeepSeek V4 Pro at 12.9%, while Vals ranked it the top open-weight system on the Vals Index, Vibe Code Bench and Terminal Bench. It also reached first place on Design Arena with a 1,360 Elo, up 27 points.

The MIT-licensed model, which has a 1 million-token context window and the same pricing as GLM-5.1, also scored 51 on Artificial Analysis's Intelligence Index and 1,524 on GDPval-AA v2, underscoring its competitiveness on longer agentic workloads. Artificial Analysis said CritPt is independently benchmarked using privately graded problems developed by Argonne and UIUC with contributions from more than 60 researchers, and even GPT-5.5 Pro solves under a third of the questions. GLM-5.2 is also spreading across deployment options, including Go support and private encrypted access on Venice.

From the sources (25 posts)

@zai_org

Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits,

@zai_org

For GLM-5.2, we strengthened 1M-context training for coding agents across large-scale implementation, automated research, performance optimization, and complex debugging. The result is a long-context system that is both broad in scope and r

@zephyr_z9

Zhipu is about to drop a banger

@zephyr_z9

HOLY SHIT!!

@victormustar

GLM-5.2 is available on Hugging Face 🔥 It's an important day for open source AI: opus-class frontier intelligence, 1M context, agentic-first by design. --> The future of AI and humanity is open

@clementdelangue

RT @Zai_org: Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long…

@zephyr_z9

Zhipu did an absolutely crazy drop

@teortaxestex

> We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. We also improve GLM-5.2’s MTP Good demonstration that a lot could be squeezed o

@testingcatalog

ZAI 🔥: GLM-5.2 is now available on huggingface! > It comes with a 1M context window and 2 levels of reasoning effort, max and high. MIT license and same pricing as GLM-5.1. > GLM-5.2 scores 46.2% on DeepSWE, the SOTA score among op

@lmsysorg

🎉 Meet GLM-5.2 from @Zai_org, the new flagship for long-horizon tasks built on a solid 1M-token context. Day-0 support is now live in SGLang! ✅ Advanced coding capability: 81.0 on Terminal-Bench 2.1 (vs 62.0 for GLM-5.1) ✅ Solid 1M contex

@calebfahlgren

RT @Zai_org: Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long…

@zai_org

RT @ProximalHQ: GLM 5.2 ranks #3 on FrontierSWE. It is only behind Fable 5 and Opus 4.8, and it outperforms GPT-5.5. This is the first mod…

@zai_org

RT @Designarena: BREAKING: GLM-5.2 is now 1st on Design Arena. With an Elo of 1360, GLM-5.2 has jumped ahead of the now unavailable Claude…

@teortaxestex

GLM 5.2 is apparently the strongest Chinese model, and it'll be open this is quite a jump on CritPt. They have *finally* started to build general reasoners. I think it'll be very competitive on ARC-AGI 2 and WeirdML, too

@zephyr_z9

RT @teortaxesTex: GLM 5.2 is apparently the strongest Chinese model, and it'll be open this is quite a jump on CritPt. They have *finally*…

@zephyr_z9

Zhipu has the best post-training stack in China

@arena

GLM-5.2 (Max) by @Zai_org ranks #10 on the new Agent Arena leaderboard, closely matching Claude-Opus-4.8 (non-thinking) and is the #1 open model by a wide margin! In Agent Arena, we measure models on millions of real-world, long-horizon ag

@victormustar

huge

@adinayakup

GLM 5.2 is here 🔥 ✨ 753B ( smaller than you expect? 👀) ✨ 1M context ✨ MIT license ✨ GLM IndexShare: reuses the indexer across layers, 2.9x fewer FLOPs/token at 1M ✨ AIME 2026: 99.2 (beats GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.8 ✨ vLL

@sentdex

Zai was gracious enough to give me a key to test out GLM 5.2. I used it on a few simple tasks and quickly realized this model is on another level. I committed to using GLM 5.2 solely for the weekend and yesterday on everything from simple

@thezachmueller

RT @Zai_org: Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long…

@zephyr_z9

@ChrissGPT Here

@kimmonismus

Lets go, GLM-5.2 released as Open Weights model. tl;dr -1M context window -MIT-licensed open weights -Stronger long-horizon coding agents -Two reasoning modes: max and high -Same API pricing as GLM-5.1 Zai says GLM-5.2 was trained specif

@andrewcurran_

Extremely powerful and MIT open-source. Mutuals who had early access are giving it great reviews.

@nrehiew_

It looks like GLM 5.2 is the second best long horizon model available??

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive