Z.ai's GLM-5.2 Scores 64% on Vibe Code Bench, Only Open Weight Model Above 60%
Benchmark provider Vals AI put Z.ai's GLM-5.2, an AI model with publicly released weights, at 64% on Vibe Code Bench v1.1, making it the only open-weight model above 60% on the test of whether systems can build web applications from scratch. No other open-weight model on the leaderboard reached 50%, leaving GLM-5.2 14 percentage points ahead of the next open-weight entry; it ranked eighth overall, ahead of GPT 5.3 Codex, GPT 5.2 and Gemini 3.5 Flash.
The score was more than double GLM-5.1's 31.5% in April and up from 3.1% for GLM 4.6 last September, based on Vals AI's figures. Earlier this week, Artificial Analysis's new AA-Briefcase benchmark for long-horizon knowledge work put GLM-5.2 at 1,266 Elo, about 90 points behind Claude Opus 4.8, while estimating a cost of roughly $2.40 per task; it had also ranked the model first among open-weight systems on its Intelligence Index with a score of 51.
From the sources (25 posts)
@teortaxestexan old cloud backup GLM-130B was the first model that made me suspect that the Chinese will be carrying the project of open source AGI on behalf of all us barbarians. DeepSeek confirmed it and provided a paradigm. It's been a damn good ride
@nrehiew_Even with compaction, this requires the team to serve an effective 1M context. The main architectural change is sharing the same top k indices across 4 layers. This reduces the O(tokens) compute even more. Worth noting that the FLOPs curv
@awnihannunRT @pcuenq: GLM 5.2 has just been released 🔥 Here it's already running with MLX on two Mac Studios (M3 Ultra). This is comparable to the…
@artificialanlysZ ai’s GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index scoring 51 and it sits on the Pareto frontier of Intelligence vs Cost per Task @Zai_org’s GLM-5.2 is the same size as GLM-5.1 (744B total /
@artificialanlysGLM-5.2 leads all open weights models on GDPval-AA v2, our primary metric for real-world agentic performance. At 1524 it places ahead of MiniMax-M3 (1418) and DeepSeek V4 Pro (max, 1328), and is effectively level with GPT-5.5 (xhigh, 1514).
@artificialanlysGLM-5.2 scores 4 on the AA-Omniscience Index, up from GLM-5.1 (2). The gain comes from both higher accuracy (25.1% vs 24.2%) and a lower hallucination rate (28.1% vs 29.4%), with attempt rate flat at 47%
@artificialanlysGLM-5.2 uses 43k output tokens per Intelligence Index task, of which 37k is reasoning. This is up from GLM-5.1 (26k) and higher than open weights peers MiniMax-M3 (24k) and Kimi K2.6 (35k), placing it among the less token-efficient open wei
@artificialanlysBreakdown of the individual evaluations in the Artificial Analysis Intelligence Index v4.1
@zephyr_z9LFG!!! 51 at AA @teortaxesTex
@testingcatalogZAI 🔥: GLM-5.2 by @Zai_org scored 51 point on Artificial Analysis Intelligence Index and got placed on the 4th spot! This made GLM-5.2 a new SOTA open-weight model. Besides that, GLM-5.2 got ranked second on Frontend Code Arena, after curr
@scaling01RT @ZixuanLi_: Finally, Artificial Analysis Intelligence Index concludes the GLM-5.2 release.
@wesrothThe Artificial Analysis Intelligence Index v4.1 is now live, shifting its evaluation toward harder, longer, and more realistic agentic workloads. The update replaces older benchmarks with: • Terminal-Bench 2.1 for complex computer tasks •
@askalphaxivRT @askalphaxiv: Introducing GLM-5.2 for understanding research papers 🚀 Highlight any section of a paper to ask questions and “@” othe…
@theahmadosmanGLM 5.2
@victormustarRT @sam_paech: GLM-5.2 is an extremely strong open weights model. Kudos to the @Zai_org team.
@jeremyphowardRT @xeophon: 51 is the same score as GPT-5.4 xhigh, btw. The model was released 3 months ago and frontier at the time
@adinayakupReally cool to see the GLM 5.2 blog on @huggingface 🔥
@jeremyphowardRT @hallerite: GLM5.2 brings back the critic. It was just a matter of time until we people would realize that group-based variance reducti…
@jeremyphowardRT @ml_angelopoulos: Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend co…
@valsaiGLM 5.2 had major improvements from its predecessor, GLM 5.1, with a 13% gain on the Vals Index and a 31% jump on Vibe Code Bench It also took the #1 spot on Terminal Bench 2.1 from Kimi K2.7 Code, an 11% improvement over GLM 5.1 https://t
@valsaiGLM 5.2 is the new open-weight SOTA on the Vals Index, Vibe Code Bench and Terminal Bench! It is also #5 across all models, and right on the heels of Opus 4.7 - released only two months ago
@artificialanlysContext on the result: CritPt is hard. It focuses on frontier physics problems developed by Argonne and UIUC through contributions from 60+ researchers globally, with the answer key and grading kept private. Models are independently benchma
@artificialanlysA standout number in Z ai’s GLM-5.2 launch is CritPt, a benchmark of unpublished research-level physics problems where it ties with Claude Opus 4.8 and is well above other open weights models Key takeaways: ➤ @Zai_org ’s GLM-5.2 (max reas
@zai_orgRT @opencode: GLM-5.2 now available in Go text · 1M context · same pricing as 5.1
@mtsliveSITUATION EXPLAINED: Z AI, a Chinese lab, just released a near-frontier model. • Competitive with everything except the very latest Opus, Fable, and GPT models • Pricing: $1 to $4 per million tokens, Haiku-tier • First open source Chinese