Command Palette
Search for a command to run...

Z.ai's GLM-5.2 Leads Open-Weight Models, Ranks No. 3 on GDPval-AA

aiai-modelingai-research-evalsai-open-modelsai-model-releases 105 posts · 47 accounts

AI lab Z.ai's GLM-5.2, an AI model released with public weights, ranked third overall on Artificial Analysis's GDPval-AA benchmark, scoring 1,524 Elo across 1,999 matches and averaging about 31 turns per task. GDPval-AA measures long-horizon, economically valuable knowledge-work tasks such as a retail supervisor's daily task list, an IEC emergency-stop circuit schematic and a music-video moodboard. Artificial Analysis placed GLM-5.2 behind Anthropic's Claude Fable 5 and Claude Opus 4.8, roughly level with GPT-5.5 and more than 100 Elo points ahead of the next open model, MiniMax-M3, at 1,408.

The pattern held on AA-Briefcase, Artificial Analysis's benchmark for multiweek projects such as financial models, board presentations and design mock-ups, where GLM-5.2 again topped open-weight models and trailed only Claude Fable 5. Artificial Analysis estimated GLM-5.2 max came within 90 Elo points of Claude Opus 4.8 while costing 65% less at $2.40 per task. In a separate test, AlphaXiv called GLM-5.2 the first open-weight model it had tried in an autoresearch pipeline that could handle real research tasks across two 8xH100 nodes.

From the sources (25 posts)

@teortaxestex

what the hell do they expect from the next Qwen-Max? GLM 5.2 wipes the floor with 3.7 (makes sense tbh! 1.5 versions ahead!) Alibaba would likely have to solidly match or exceed Opus 4.8. Or do they mean something different from "the compan

@zai_org

RT @FireworksAI_HQ: "...at least as good as Opus 4.8 and GPT 5.5."

@thezachmueller

RT @_xjdr: after spending a ton of time with GLM5.2 today in order to add it to noumena, i have to say i am very impressed. if it keeps thi…

@thezachmueller

RT @_xjdr: To continue the celebration, we have added GLM 5.2 support to ncode and the noumena platform and are making it free to use for t…

@teortaxestex

In practice GLM 5.2 as part of the Zhipu subscription product "has vision", which I have just now learned. It seems they resort to calling GLM-4.5V via MCP. It's not a big deal tbh. Their business is selling coding plans. They can afford f

@clementdelangue

RT @elliotarledge: KernelBench-Hard and KernelBench-Mega results are in. Reasoning traces are open. Thanks to @calebfahlgren for showing me…

@clementdelangue

RT @PatrickToulme: I ran GLM 5.2 with OpenCode harness against Claude Opus this week deployed locally. Bottom line: It is a real frontier…

@yuchenj_uw

After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by side with Opus 4.8, and sometimes I even preferred GLM-5.2’s results. OSS LLMs are impressive, especially given how

@brianroemmele

Like I said Open Source Anthropic Mythos class AI in GLM-5.2! We see the same. Time to pick a different bogeyman for Anthropic, this is now in everybody’s hands.

@clementdelangue

RT @elliotarledge: post coming shortly

@clementdelangue

RT @elliotarledge: I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on…

@clementdelangue

RT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…

@aravsrinivas

RT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…

@teortaxestex

by the way, "DNF" is not "it categorically cannot write the kernel", it's "Elliot got rate limited". GLM 5.2 is at the frontier in kernel engineering, simple as.

@josephjacks_

RT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…

@hrishioa

RT @hrishioa: This is a watershed moment. GLM-5.2 solidly beat Opus 4.8 and human participants in our backend take-home, making the whole…

@wesroth

RT @WesRoth: GLM-5.2 became the highest-ranked open-weight model across the Vals Index, Vibe Code Bench, and Terminal-Bench 2.1. It ranks…

@wesroth

RT @WesRoth: GLM-5.2 Max is now the leading open-weight model on the Artificial Analysis Intelligence Index v4.1. It scored 51, placing it…

@teortaxestex

As impressive as GLM 5.2 is, at the end of the day it's ≈5-10X more expensive than DeepSeek V4 for the same-sized session, and they can't serve the demand. If V4.1 is noticeably but not crushingly worse, it takes a lot of marginal customers

@eliebakouch

french president emmanuel macron coming out of stealth and open sourcing a megatron lm fork is unexpected

@teortaxestex

RT @slime_framework: Thanks for the support! A small note: slime has supported not only OPD, but the full RL + OPD post-training workflow…

@teortaxestex

GLM 5.2 is one *of the* greatest gap reductions ever, but I think it is *the* greatest show of benchmark solidity from an open model claiming SoTA ever. Normally, you have some variety of the bad old Qwen pattern: headline benchmarks are So

@ben_burtenshaw

RT @elliotarledge: I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on…

@jeremyphoward

RT @didier_lopes: Incredible how Z. ai literally has their RL infrastructure open source. The entire OPD post-training of GLM-5.2 took on…

@arena

RT @hqmank: Agent Arena AI lab ranking update: Zai moved to #3 after releasing GLM 5.2, ahead of Google. The current top 10 now includes…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive