Z.ai's GLM-5.2 Ranks No. 3 on FrontierSWE, With 44% DeepSWE Pass Rate
GLM-5.2, the latest model from Chinese AI lab Z.ai released with public weights, posted a 44% first-attempt pass rate on DeepSWE, a software-engineering benchmark, and ranked third on FrontierSWE, which measures long-horizon engineering tasks. FrontierSWE placed it behind Anthropic's Claude Fable 5 and Opus 4.8 but ahead of GPT-5.5, while Code Arena's frontend leaderboard put GLM-5.2 Max second behind Fable 5.
Performance outside coding has looked more uneven. Matharena measured only a 1.9% expected-performance gain over GLM-5.1, and several posts described heavier output-token use that can make GLM-5.2 slower and more expensive in practice than Claude Opus 4.8 or GPT-5.5 on medium settings despite cheaper token prices. Claims that the broader GLM-5 line was trained on Huawei Ascend chips remain unverified; separate posts linked Ascend use to post-training or reinforcement-learning rollouts rather than base-model training.
From the sources (25 posts)
@artificialanlys@jeremyphoward @Zai_org GLM-5.2 is between GPT-5.5 and Opus 4.8 in our new agentic knowledge work eval that we released just today. Very impressive
@nandodfRT @rasbt: Just caught up with the recent GLM-5.2 release. The best open-weight model today. Architecture-wise, it's build on the GLM-5 an…
@_arohan_RT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…
@andrewcurran_Rave reviews all over my timeline for GLM 5.2 the last 48 hours, open source is having an exceptional month.
@clementdelangue“Cost per task varies by ~800x across models tested: Claude Fable 5 leads the benchmark but costs more than $31 per task on average, compared to ~$0.04 for DeepSeek V4 Flash (max). The strongest price/performance options are open weights mo
@chris_j_paxtonRT @jietang: GLM-5.2 is Fully Open, Frontier Intelligence Belongs to Everyone Today, the sudden restriction of certain frontier models is…
@clementdelangueRT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…
@ollamaRT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…
@clementdelangueRT @AndrewCurran_: Rave reviews all over my timeline for GLM 5.2 the last 48 hours, open source is having an exceptional month.
@zai_orgRT @ArtificialAnlys: Announcing AA-Briefcase, the benchmark for the next era of agentic knowledge work AA-Briefcase is our new benchmark f…
@andersonbcdefgRT @teortaxesTex: ok, here's how GLM 5.2 performs on a bench it definitely didn't see, and where GLM 5.1 scored 0.0%. Closer to Opus 4.8 th…
@matvellosoAll day using GLM 5.2. Didn't miss much. First open model that passes the bar as a daily driver. Things are not going to be the same. Damn, now I want to buy some serious hardware.
@wesrothRT @WesRoth: GLM-5.2 Max reached second place in the Code Arena Frontend leaderboard, behind only the currently unavailable Claude Fable 5.…
@quixiaiRT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…
@zai_orgRT @ZixuanLi_: GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Res…
@zai_orgLong-horizon is more than a concept. It should live in real-world scenarios, empowering AI builders to solve the problems that matter. And more scenarios are on the way.
@teortaxestexZhipu post-training team is very very good
@zai_orgRT @CunxiangWang: GLM-5.2 is not only stronger on benchmarks, but also much better in real app development scenarios — iOS, Android, WeChat…
@teortaxestexGLM results are insanely high If this is "distillation", it's doing better than any previous attempt Eg ProofBench should be amenable both to distillation and RLVR, yet V4, Kimi, Mimo, Grok lol are all hopeless. Zhipu will obviously aim to
@teortaxestexRT @banteg: not a bad result from glm-5.2, it found this after gpt-5.5 xhigh
@quixiaiRT @AlexFinn: I can't believe this is real I have GLM 5.2 running 100% locally on my Mac Studio. 2 bit quant. The results I'm getting are…
@vipulvedSpent the last 3 hours playing with GLM-5.2. It’s truly a stunning model. What a time to be alive.
@ollamaRT @matvelloso: All day using GLM 5.2. Didn't miss much. First open model that passes the bar as a daily driver. Things are not going to be…
@matansfRT @alphatozeta8148: @baseten model performance team is absolutely cracked. @Zai_org GLM 5.2 is now 4x faster running at full 1M context!…
@teortaxestexInteresting that in "GLM 5.2 is on par with X model", X is so widely distributed over different capabilities, and it's not like the gap is just greater for the difficult ones. It plainly can't do some parlor tricks Sonnet 3.5 pulled off. In