Z.ai's Open-Weight GLM-5.2 Completes Research Workflow, Scores 44% on DeepSWE
Chinese AI lab Z.ai's GLM-5.2 was the first open-weight model AlphaXiv had tested in its autoresearch pipeline that it judged capable of real research tasks. In a demonstration across two 8xH100 nodes, the model ran asynchronous and colocated reinforcement-learning training, resolved setup issues, tracked runs to completion and produced a comparison of throughput and reward stability. The result follows GLM-5.2's recent rollout on Together Chat and Fireworks.
Separate results on DeepSWE v1.1, a software-engineering benchmark, put GLM-5.2 Max at 44%, with average usage of $3.92 per task, 78,000 output tokens and 129 agent steps. That beat the tested Gemini 3.5 Flash Medium at 37% and Gemini 3.1 Pro High at 12%, but trailed GPT-5.4 xhigh at 52%. AlphaXiv also noted that GLM-5.2 lacks image understanding, so instead of reading Weights & Biases, or WandB, charts directly it analyzed raw numbers by writing NumPy code, which the group said could become a limitation on larger sweeps or ablations.
From the sources (25 posts)
@clementdelangueRT @gneubig: OK, I tried GLM-5.2 and this is a good model. Probably the first model good enough to eschew closed models from your workflow…
@clementdelangueRT @matvelloso: All day using GLM 5.2. Didn't miss much. First open model that passes the bar as a daily driver. Things are not going to be…
@valsaiGLM 5.2 is the only open-weight model to break 60% on Vibe Code Bench v1.1, our test of whether models can build web applications from scratch It scores 64%, and no other open-weight model on the board reaches even 50%. That puts it 14 per
@valsaiThe progress across generations is exceptional- GLM 5.2 scored 64% which is more than double what GLM 5.1 reached in April (31.5%), and up from 3.1% for GLM 4.6 last September. That is a 60.9 point gain over four releases in roughly nine mo
@valsai@Zai_org is now competing with frontier models. On the full leaderboard, GLM 5.2 ranks 8th, ahead of GPT 5.3 Codex, GPT 5.2, Gemini 3.5 Flash, and several Opus 4.6 and Sonnet 4.6 configurations. An open-weight model now sits among frontier
@altryneThe timeline agrees.. GLM 5.2 vibes are immaculate! Congrats @Zai_org for this awesome model! I asked this model to create a unique episode page for itself, and whoah it delivered! Look at this beauty 😍 It works
@teortaxestexwhat the hell do they expect from the next Qwen-Max? GLM 5.2 wipes the floor with 3.7 (makes sense tbh! 1.5 versions ahead!) Alibaba would likely have to solidly match or exceed Opus 4.8. Or do they mean something different from "the compan
@zai_orgRT @FireworksAI_HQ: "...at least as good as Opus 4.8 and GPT 5.5."
@thezachmuellerRT @_xjdr: after spending a ton of time with GLM5.2 today in order to add it to noumena, i have to say i am very impressed. if it keeps thi…
@thezachmuellerRT @_xjdr: To continue the celebration, we have added GLM 5.2 support to ncode and the noumena platform and are making it free to use for t…
@teortaxestexIn practice GLM 5.2 as part of the Zhipu subscription product "has vision", which I have just now learned. It seems they resort to calling GLM-4.5V via MCP. It's not a big deal tbh. Their business is selling coding plans. They can afford f
@clementdelangueRT @elliotarledge: KernelBench-Hard and KernelBench-Mega results are in. Reasoning traces are open. Thanks to @calebfahlgren for showing me…
@clementdelangueRT @PatrickToulme: I ran GLM 5.2 with OpenCode harness against Claude Opus this week deployed locally. Bottom line: It is a real frontier…
@yuchenj_uwAfter using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by side with Opus 4.8, and sometimes I even preferred GLM-5.2’s results. OSS LLMs are impressive, especially given how
@brianroemmeleLike I said Open Source Anthropic Mythos class AI in GLM-5.2! We see the same. Time to pick a different bogeyman for Anthropic, this is now in everybody’s hands.
@clementdelangueRT @elliotarledge: post coming shortly
@clementdelangueRT @elliotarledge: I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on…
@clementdelangueRT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…
@aravsrinivasRT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…
@teortaxestexby the way, "DNF" is not "it categorically cannot write the kernel", it's "Elliot got rate limited". GLM 5.2 is at the frontier in kernel engineering, simple as.
@josephjacks_RT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…
@hrishioaRT @hrishioa: This is a watershed moment. GLM-5.2 solidly beat Opus 4.8 and human participants in our backend take-home, making the whole…
@wesrothRT @WesRoth: GLM-5.2 became the highest-ranked open-weight model across the Vals Index, Vibe Code Bench, and Terminal-Bench 2.1. It ranks…
@wesrothRT @WesRoth: GLM-5.2 Max is now the leading open-weight model on the Artificial Analysis Intelligence Index v4.1. It scored 51, placing it…
@teortaxestexAs impressive as GLM 5.2 is, at the end of the day it's ≈5-10X more expensive than DeepSeek V4 for the same-sized session, and they can't serve the demand. If V4.1 is noticeably but not crushingly worse, it takes a lot of marginal customers