Command Palette
Search for a command to run...

Z.ai's Open-Weight GLM-5.2 Trails Claude Opus 4.8 by 90 Elo on AA-Briefcase, Costs $2.40 a Task

aiai-modelingai-open-modelsai-model-releasesai-research-evals 134 posts · 57 accounts

Z.ai's GLM-5.2, an MIT-licensed open-weight AI model with a 1 million-token context window, scored 1,266 on Artificial Analysis's newly launched AA-Briefcase benchmark for long-horizon knowledge-work projects. That left it 90 Elo behind Claude Opus 4.8 at 1,356 and below Claude Fable 5 at 1,587, while costing about $2.40 per task versus $10.40 for Opus 4.8 and more than $31 for Fable 5. Artificial Analysis said GLM-5.2 falls between GPT-5.5 and Opus 4.8 on the test.

Z.ai separately said GLM-5.2 completed 48 of 70 trials on an internal mobile-app development benchmark, up from 21 of 70 for GLM-5.1 and below Claude Fable 5's 56 of 70. Earlier results had already put the model at 51 on Artificial Analysis's Intelligence Index, while Vals ranked it first on the Vals Index, Harvey's Legal Agent Benchmark, Finance Agent v2, ProofBench and Vibe Code Bench.

From the sources (25 posts)

@zai_org

Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long-horizon capabilities with a 1M context window - Two levels of reasoning effort: GLM-5.2 (max) pushes the limits,

@zai_org

For GLM-5.2, we strengthened 1M-context training for coding agents across large-scale implementation, automated research, performance optimization, and complex debugging. The result is a long-context system that is both broad in scope and r

@zephyr_z9

Zhipu is about to drop a banger

@zephyr_z9

HOLY SHIT!!

@victormustar

GLM-5.2 is available on Hugging Face 🔥 It's an important day for open source AI: opus-class frontier intelligence, 1M context, agentic-first by design. --> The future of AI and humanity is open

@clementdelangue

RT @Zai_org: Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long…

@zephyr_z9

Zhipu did an absolutely crazy drop

@teortaxestex

> We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. We also improve GLM-5.2’s MTP Good demonstration that a lot could be squeezed o

@testingcatalog

ZAI 🔥: GLM-5.2 is now available on huggingface! > It comes with a 1M context window and 2 levels of reasoning effort, max and high. MIT license and same pricing as GLM-5.1. > GLM-5.2 scores 46.2% on DeepSWE, the SOTA score among op

@lmsysorg

🎉 Meet GLM-5.2 from @Zai_org, the new flagship for long-horizon tasks built on a solid 1M-token context. Day-0 support is now live in SGLang! ✅ Advanced coding capability: 81.0 on Terminal-Bench 2.1 (vs 62.0 for GLM-5.1) ✅ Solid 1M contex

@calebfahlgren

RT @Zai_org: Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long…

@zai_org

RT @ProximalHQ: GLM 5.2 ranks #3 on FrontierSWE. It is only behind Fable 5 and Opus 4.8, and it outperforms GPT-5.5. This is the first mod…

@zai_org

RT @Designarena: BREAKING: GLM-5.2 is now 1st on Design Arena. With an Elo of 1360, GLM-5.2 has jumped ahead of the now unavailable Claude…

@teortaxestex

GLM 5.2 is apparently the strongest Chinese model, and it'll be open this is quite a jump on CritPt. They have *finally* started to build general reasoners. I think it'll be very competitive on ARC-AGI 2 and WeirdML, too

@zephyr_z9

RT @teortaxesTex: GLM 5.2 is apparently the strongest Chinese model, and it'll be open this is quite a jump on CritPt. They have *finally*…

@zephyr_z9

Zhipu has the best post-training stack in China

@arena

GLM-5.2 (Max) by @Zai_org ranks #10 on the new Agent Arena leaderboard, closely matching Claude-Opus-4.8 (non-thinking) and is the #1 open model by a wide margin! In Agent Arena, we measure models on millions of real-world, long-horizon ag

@victormustar

huge

@adinayakup

GLM 5.2 is here 🔥 ✨ 753B ( smaller than you expect? 👀) ✨ 1M context ✨ MIT license ✨ GLM IndexShare: reuses the indexer across layers, 2.9x fewer FLOPs/token at 1M ✨ AIME 2026: 99.2 (beats GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.8 ✨ vLL

@sentdex

Zai was gracious enough to give me a key to test out GLM 5.2. I used it on a few simple tasks and quickly realized this model is on another level. I committed to using GLM 5.2 solely for the weekend and yesterday on everything from simple

@thezachmueller

RT @Zai_org: Introducing GLM-5.2: Frontier Intelligence, Open Weights - Significant improvements in coding and agentic tasks - Strong long…

@zephyr_z9

@ChrissGPT Here

@kimmonismus

Lets go, GLM-5.2 released as Open Weights model. tl;dr -1M context window -MIT-licensed open weights -Stronger long-horizon coding agents -Two reasoning modes: max and high -Same API pricing as GLM-5.1 Zai says GLM-5.2 was trained specif

@andrewcurran_

Extremely powerful and MIT open-source. Mutuals who had early access are giving it great reviews.

@nrehiew_

It looks like GLM 5.2 is the second best long horizon model available??

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive