Command Palette
Search for a command to run...

Z.ai's GLM-5.2 Scores 64% on Vibe Code Bench, Only Open Weight Model Above 60%

aiai-modelingai-research-evalsai-open-models 75 posts · 40 accounts

Benchmark provider Vals AI put Z.ai's GLM-5.2, an AI model with publicly released weights, at 64% on Vibe Code Bench v1.1, making it the only open-weight model above 60% on the test of whether systems can build web applications from scratch. No other open-weight model on the leaderboard reached 50%, leaving GLM-5.2 14 percentage points ahead of the next open-weight entry; it ranked eighth overall, ahead of GPT 5.3 Codex, GPT 5.2 and Gemini 3.5 Flash.

The score was more than double GLM-5.1's 31.5% in April and up from 3.1% for GLM 4.6 last September, based on Vals AI's figures. Earlier this week, Artificial Analysis's new AA-Briefcase benchmark for long-horizon knowledge work put GLM-5.2 at 1,266 Elo, about 90 points behind Claude Opus 4.8, while estimating a cost of roughly $2.40 per task; it had also ranked the model first among open-weight systems on its Intelligence Index with a score of 51.

From the sources (25 posts)

@teortaxestex

an old cloud backup GLM-130B was the first model that made me suspect that the Chinese will be carrying the project of open source AGI on behalf of all us barbarians. DeepSeek confirmed it and provided a paradigm. It's been a damn good ride

@nrehiew_

Even with compaction, this requires the team to serve an effective 1M context. The main architectural change is sharing the same top k indices across 4 layers. This reduces the O(tokens) compute even more. Worth noting that the FLOPs curv

@awnihannun

RT @pcuenq: GLM 5.2 has just been released 🔥 Here it's already running with MLX on two Mac Studios (M3 Ultra). This is comparable to the…

@artificialanlys

Z ai’s GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index scoring 51 and it sits on the Pareto frontier of Intelligence vs Cost per Task @Zai_org’s GLM-5.2 is the same size as GLM-5.1 (744B total /

@artificialanlys

GLM-5.2 leads all open weights models on GDPval-AA v2, our primary metric for real-world agentic performance. At 1524 it places ahead of MiniMax-M3 (1418) and DeepSeek V4 Pro (max, 1328), and is effectively level with GPT-5.5 (xhigh, 1514).

@artificialanlys

GLM-5.2 scores 4 on the AA-Omniscience Index, up from GLM-5.1 (2). The gain comes from both higher accuracy (25.1% vs 24.2%) and a lower hallucination rate (28.1% vs 29.4%), with attempt rate flat at 47%

@artificialanlys

GLM-5.2 uses 43k output tokens per Intelligence Index task, of which 37k is reasoning. This is up from GLM-5.1 (26k) and higher than open weights peers MiniMax-M3 (24k) and Kimi K2.6 (35k), placing it among the less token-efficient open wei

@artificialanlys

Breakdown of the individual evaluations in the Artificial Analysis Intelligence Index v4.1

@zephyr_z9

LFG!!! 51 at AA @teortaxesTex

@testingcatalog

ZAI 🔥: GLM-5.2 by @Zai_org scored 51 point on Artificial Analysis Intelligence Index and got placed on the 4th spot! This made GLM-5.2 a new SOTA open-weight model. Besides that, GLM-5.2 got ranked second on Frontend Code Arena, after curr

@scaling01

RT @ZixuanLi_: Finally, Artificial Analysis Intelligence Index concludes the GLM-5.2 release.

@wesroth

The Artificial Analysis Intelligence Index v4.1 is now live, shifting its evaluation toward harder, longer, and more realistic agentic workloads. The update replaces older benchmarks with: • Terminal-Bench 2.1 for complex computer tasks •

@askalphaxiv

RT @askalphaxiv: Introducing GLM-5.2 for understanding research papers 🚀 Highlight any section of a paper to ask questions and “@” othe…

@theahmadosman

GLM 5.2

@victormustar

RT @sam_paech: GLM-5.2 is an extremely strong open weights model. Kudos to the @Zai_org team.

@jeremyphoward

RT @xeophon: 51 is the same score as GPT-5.4 xhigh, btw. The model was released 3 months ago and frontier at the time

@adinayakup

Really cool to see the GLM 5.2 blog on @huggingface 🔥

@jeremyphoward

RT @hallerite: GLM5.2 brings back the critic. It was just a matter of time until we people would realize that group-based variance reducti…

@jeremyphoward

RT @ml_angelopoulos: Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend co…

@valsai

GLM 5.2 had major improvements from its predecessor, GLM 5.1, with a 13% gain on the Vals Index and a 31% jump on Vibe Code Bench It also took the #1 spot on Terminal Bench 2.1 from Kimi K2.7 Code, an 11% improvement over GLM 5.1 https://t

@valsai

GLM 5.2 is the new open-weight SOTA on the Vals Index, Vibe Code Bench and Terminal Bench! It is also #5 across all models, and right on the heels of Opus 4.7 - released only two months ago

@artificialanlys

Context on the result: CritPt is hard. It focuses on frontier physics problems developed by Argonne and UIUC through contributions from 60+ researchers globally, with the answer key and grading kept private. Models are independently benchma

@artificialanlys

A standout number in Z ai’s GLM-5.2 launch is CritPt, a benchmark of unpublished research-level physics problems where it ties with Claude Opus 4.8 and is well above other open weights models Key takeaways: ➤ @Zai_org ’s GLM-5.2 (max reas

@zai_org

RT @opencode: GLM-5.2 now available in Go text · 1M context · same pricing as 5.1

@mtslive

SITUATION EXPLAINED: Z AI, a Chinese lab, just released a near-frontier model. • Competitive with everything except the very latest Opus, Fable, and GPT models • Pricing: $1 to $4 per million tokens, Haiku-tier • First open source Chinese

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive