Command Palette
Search for a command to run...

Z.ai's Open-Weight GLM-5.2 Completes Research Workflow, Scores 44% on DeepSWE

aiai-modelingai-research-evalsai-open-modelsai-model-releases 101 posts · 46 accounts

Chinese AI lab Z.ai's GLM-5.2 was the first open-weight model AlphaXiv had tested in its autoresearch pipeline that it judged capable of real research tasks. In a demonstration across two 8xH100 nodes, the model ran asynchronous and colocated reinforcement-learning training, resolved setup issues, tracked runs to completion and produced a comparison of throughput and reward stability. The result follows GLM-5.2's recent rollout on Together Chat and Fireworks.

Separate results on DeepSWE v1.1, a software-engineering benchmark, put GLM-5.2 Max at 44%, with average usage of $3.92 per task, 78,000 output tokens and 129 agent steps. That beat the tested Gemini 3.5 Flash Medium at 37% and Gemini 3.1 Pro High at 12%, but trailed GPT-5.4 xhigh at 52%. AlphaXiv also noted that GLM-5.2 lacks image understanding, so instead of reading Weights & Biases, or WandB, charts directly it analyzed raw numbers by writing NumPy code, which the group said could become a limitation on larger sweeps or ablations.

From the sources (25 posts)

@clementdelangue

RT @gneubig: OK, I tried GLM-5.2 and this is a good model. Probably the first model good enough to eschew closed models from your workflow…

@clementdelangue

RT @matvelloso: All day using GLM 5.2. Didn't miss much. First open model that passes the bar as a daily driver. Things are not going to be…

@valsai

GLM 5.2 is the only open-weight model to break 60% on Vibe Code Bench v1.1, our test of whether models can build web applications from scratch It scores 64%, and no other open-weight model on the board reaches even 50%. That puts it 14 per

@valsai

The progress across generations is exceptional- GLM 5.2 scored 64% which is more than double what GLM 5.1 reached in April (31.5%), and up from 3.1% for GLM 4.6 last September. That is a 60.9 point gain over four releases in roughly nine mo

@valsai

@Zai_org is now competing with frontier models. On the full leaderboard, GLM 5.2 ranks 8th, ahead of GPT 5.3 Codex, GPT 5.2, Gemini 3.5 Flash, and several Opus 4.6 and Sonnet 4.6 configurations. An open-weight model now sits among frontier

@altryne

The timeline agrees.. GLM 5.2 vibes are immaculate! Congrats @Zai_org for this awesome model! I asked this model to create a unique episode page for itself, and whoah it delivered! Look at this beauty 😍 It works

@teortaxestex

what the hell do they expect from the next Qwen-Max? GLM 5.2 wipes the floor with 3.7 (makes sense tbh! 1.5 versions ahead!) Alibaba would likely have to solidly match or exceed Opus 4.8. Or do they mean something different from "the compan

@zai_org

RT @FireworksAI_HQ: "...at least as good as Opus 4.8 and GPT 5.5."

@thezachmueller

RT @_xjdr: after spending a ton of time with GLM5.2 today in order to add it to noumena, i have to say i am very impressed. if it keeps thi…

@thezachmueller

RT @_xjdr: To continue the celebration, we have added GLM 5.2 support to ncode and the noumena platform and are making it free to use for t…

@teortaxestex

In practice GLM 5.2 as part of the Zhipu subscription product "has vision", which I have just now learned. It seems they resort to calling GLM-4.5V via MCP. It's not a big deal tbh. Their business is selling coding plans. They can afford f

@clementdelangue

RT @elliotarledge: KernelBench-Hard and KernelBench-Mega results are in. Reasoning traces are open. Thanks to @calebfahlgren for showing me…

@clementdelangue

RT @PatrickToulme: I ran GLM 5.2 with OpenCode harness against Claude Opus this week deployed locally. Bottom line: It is a real frontier…

@yuchenj_uw

After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by side with Opus 4.8, and sometimes I even preferred GLM-5.2’s results. OSS LLMs are impressive, especially given how

@brianroemmele

Like I said Open Source Anthropic Mythos class AI in GLM-5.2! We see the same. Time to pick a different bogeyman for Anthropic, this is now in everybody’s hands.

@clementdelangue

RT @elliotarledge: post coming shortly

@clementdelangue

RT @elliotarledge: I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on…

@clementdelangue

RT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…

@aravsrinivas

RT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…

@teortaxestex

by the way, "DNF" is not "it categorically cannot write the kernel", it's "Elliot got rate limited". GLM 5.2 is at the frontier in kernel engineering, simple as.

@josephjacks_

RT @Yuchenj_UW: After using GLM-5.2 for a day, I’m surprised by how often it feels close to Opus 4.8/GPT-5.5 level. I compared it side by…

@hrishioa

RT @hrishioa: This is a watershed moment. GLM-5.2 solidly beat Opus 4.8 and human participants in our backend take-home, making the whole…

@wesroth

RT @WesRoth: GLM-5.2 became the highest-ranked open-weight model across the Vals Index, Vibe Code Bench, and Terminal-Bench 2.1. It ranks…

@wesroth

RT @WesRoth: GLM-5.2 Max is now the leading open-weight model on the Artificial Analysis Intelligence Index v4.1. It scored 51, placing it…

@teortaxestex

As impressive as GLM 5.2 is, at the end of the day it's ≈5-10X more expensive than DeepSeek V4 for the same-sized session, and they can't serve the demand. If V4.1 is noticeably but not crushingly worse, it takes a lot of marginal customers

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive