Z.ai's GLM-5.2 Leads Open-Weight Models, Ranks No. 3 on GDPval-AA
AI lab Z.ai's GLM-5.2, an AI model released with public weights, ranked third overall on Artificial Analysis's GDPval-AA benchmark, scoring 1,524 Elo across 1,999 matches and averaging about 31 turns per task. GDPval-AA measures long-horizon, economically valuable knowledge-work tasks such as a retail supervisor's daily task list, an IEC emergency-stop circuit schematic and a music-video moodboard. Artificial Analysis placed GLM-5.2 behind Anthropic's Claude Fable 5 and Claude Opus 4.8, roughly level with GPT-5.5 and more than 100 Elo points ahead of the next open model, MiniMax-M3, at 1,408.
The pattern held on AA-Briefcase, Artificial Analysis's benchmark for multiweek projects such as financial models, board presentations and design mock-ups, where GLM-5.2 again topped open-weight models and trailed only Claude Fable 5. Artificial Analysis estimated GLM-5.2 max came within 90 Elo points of Claude Opus 4.8 while costing 65% less at $2.40 per task. In a separate test, AlphaXiv called GLM-5.2 the first open-weight model it had tried in an autoresearch pipeline that could handle real research tasks across two 8xH100 nodes.
From the sources (25 posts)
@teortaxestexwhat the hell do they expect from the next Qwen-Max? GLM 5.2 wipes the floor with 3.7 (makes sense tbh! 1.5 versions ahead!) Alibaba would likely have to solidly match or exceed Opus 4.8. Or do they mean something different from "the compan
@zai_orgRT @FireworksAI_HQ: "...at least as good as Opus 4.8 and GPT 5.5."
@thezachmuellerRT @_xjdr: after spending a ton of time with GLM5.2 today in order to add it to noumena, i have to say i am very impressed. if it keeps thi…
@thezachmuellerRT @_xjdr: To continue the celebration, we have added GLM 5.2 support to ncode and the noumena platform and are making it free to use for t…
@teortaxestexIn practice GLM 5.2 as part of the Zhipu subscription product "has vision", which I have just now learned. It seems they resort to calling GLM-4.5V via MCP. It's not a big deal tbh. Their business is selling coding plans. They can afford f
@clementdelangueRT @elliotarledge: KernelBench-Hard and KernelBench-Mega results are in. Reasoning traces are open. Thanks to @calebfahlgren for showing me…
@clementdelangueRT @PatrickToulme: I ran GLM 5.2 with OpenCode harness against Claude Opus this week deployed locally. Bottom line: It is a real frontier…
@yuchenj_uwAfter using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by side with Opus 4.8, and sometimes I even preferred GLM-5.2’s results. OSS LLMs are impressive, especially given how
@brianroemmeleLike I said Open Source Anthropic Mythos class AI in GLM-5.2! We see the same. Time to pick a different bogeyman for Anthropic, this is now in everybody’s hands.
@clementdelangueRT @elliotarledge: post coming shortly
@clementdelangueRT @elliotarledge: I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on…
@clementdelangueRT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…
@aravsrinivasRT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…
@teortaxestexby the way, "DNF" is not "it categorically cannot write the kernel", it's "Elliot got rate limited". GLM 5.2 is at the frontier in kernel engineering, simple as.
@josephjacks_RT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…
@hrishioaRT @hrishioa: This is a watershed moment. GLM-5.2 solidly beat Opus 4.8 and human participants in our backend take-home, making the whole…
@wesrothRT @WesRoth: GLM-5.2 became the highest-ranked open-weight model across the Vals Index, Vibe Code Bench, and Terminal-Bench 2.1. It ranks…
@wesrothRT @WesRoth: GLM-5.2 Max is now the leading open-weight model on the Artificial Analysis Intelligence Index v4.1. It scored 51, placing it…
@teortaxestexAs impressive as GLM 5.2 is, at the end of the day it's ≈5-10X more expensive than DeepSeek V4 for the same-sized session, and they can't serve the demand. If V4.1 is noticeably but not crushingly worse, it takes a lot of marginal customers
@eliebakouchfrench president emmanuel macron coming out of stealth and open sourcing a megatron lm fork is unexpected
@teortaxestexRT @slime_framework: Thanks for the support! A small note: slime has supported not only OPD, but the full RL + OPD post-training workflow…
@teortaxestexGLM 5.2 is one *of the* greatest gap reductions ever, but I think it is *the* greatest show of benchmark solidity from an open model claiming SoTA ever. Normally, you have some variety of the bad old Qwen pattern: headline benchmarks are So
@ben_burtenshawRT @elliotarledge: I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on…
@jeremyphowardRT @didier_lopes: Incredible how Z. ai literally has their RL infrastructure open source. The entire OPD post-training of GLM-5.2 took on…
@arenaRT @hqmank: Agent Arena AI lab ranking update: Zai moved to #3 after releasing GLM 5.2, ahead of Google. The current top 10 now includes…