Command Palette
Search for a command to run...

Z.ai's GLM-5.2 Scores 64% on Vibe Code Bench, Only Open-Weight Model Above 60%

aiai-modelingai-research-evalsai-open-models 81 posts · 42 accounts

Benchmark firm Vals AI ranked Chinese AI lab Z.ai's GLM-5.2, an AI model released with public weights, as the highest open-weight system across the Vals Index, Vibe Code Bench and Terminal Bench 2.1. On Vibe Code Bench v1.1, a test of whether models can build web applications from scratch, it scored 64%, making it the only open-weight model above 60% and 14 percentage points ahead of the next open-weight entry.

Separate KernelBench-Mega results, which test models writing GPU megakernels from scratch on Nvidia's RTX PRO 6000, H100 and B200 chips, put Claude Opus 4.8 from Anthropic first overall and GLM-5.2 first among open-weight models. Artificial Analysis had earlier ranked GLM-5.2 first among open-weight systems on its Intelligence Index with a score of 51, and developer tools including ncode and the Noumena platform have added support.

From the sources (25 posts)

@askalphaxiv

RT @askalphaxiv: Introducing GLM-5.2 for understanding research papers 🚀 Highlight any section of a paper to ask questions and “@” othe…

@theahmadosman

GLM 5.2

@victormustar

RT @sam_paech: GLM-5.2 is an extremely strong open weights model. Kudos to the @Zai_org team.

@jeremyphoward

RT @xeophon: 51 is the same score as GPT-5.4 xhigh, btw. The model was released 3 months ago and frontier at the time

@adinayakup

Really cool to see the GLM 5.2 blog on @huggingface 🔥

@jeremyphoward

RT @hallerite: GLM5.2 brings back the critic. It was just a matter of time until we people would realize that group-based variance reducti…

@jeremyphoward

RT @ml_angelopoulos: Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend co…

@valsai

GLM 5.2 had major improvements from its predecessor, GLM 5.1, with a 13% gain on the Vals Index and a 31% jump on Vibe Code Bench It also took the #1 spot on Terminal Bench 2.1 from Kimi K2.7 Code, an 11% improvement over GLM 5.1 https://t

@valsai

GLM 5.2 is the new open-weight SOTA on the Vals Index, Vibe Code Bench and Terminal Bench! It is also #5 across all models, and right on the heels of Opus 4.7 - released only two months ago

@artificialanlys

Context on the result: CritPt is hard. It focuses on frontier physics problems developed by Argonne and UIUC through contributions from 60+ researchers globally, with the answer key and grading kept private. Models are independently benchma

@artificialanlys

A standout number in Z ai’s GLM-5.2 launch is CritPt, a benchmark of unpublished research-level physics problems where it ties with Claude Opus 4.8 and is well above other open weights models Key takeaways: ➤ @Zai_org ’s GLM-5.2 (max reas

@zai_org

RT @opencode: GLM-5.2 now available in Go text · 1M context · same pricing as 5.1

@mtslive

SITUATION EXPLAINED: Z AI, a Chinese lab, just released a near-frontier model. • Competitive with everything except the very latest Opus, Fable, and GPT models • Pricing: $1 to $4 per million tokens, Haiku-tier • First open source Chinese

@zephyr_z9

Massive I can't believe that a 700B model can hit this

@hosseeb

RT @Shaughnessy119: I agree with @hosseeb, the gap between Open and Closed source is not insane GLM 5.2 is 90% cheaper vs Fable 5, was rel…

@wesroth

GLM-5.2 reached first place on Design Arena with an Elo score of 1360. The model gained 27 Elo points and climbed four positions, moving ahead of the now-unavailable Claude Fable 5.

@mtslive

SITUATION ANALYSIS: New Model From Z AI? Today, Chinese startup Z AI just released GLM-5.2, an open-source model closer to the frontier than any other in history. I know you’re probably used to hearing about benchmaxxed open-source slop mo

@chamath

It may be overfit so it’s a bit early to declare any sort of victory but it is really good to see open-weight models catch up to the closed source labs so quickly after their latest releases. With recursive RL, the delta in time between a

@blackhc

RT @ProximalHQ: GLM 5.2 ranks #3 on FrontierSWE. It is only behind Fable 5 and Opus 4.8, and it outperforms GPT-5.5. This is the first mod…

@tftc21

China just shipped a 744B parameter frontier AI model. MIT license, runs on consumer hardware, no NVIDIA required, 1M token context window. Dropped one day after the US suspended access to Fable 5. Open-source AI is freedom tech on th

@kimmonismus

GLM-5.2 max is currently the third best model available, across both open and proprietary options. And that's fantastic. Open source is fundamentally important and must continue to hold a strong position so that everyone has open alternati

@andykonwinski

RT @ml_angelopoulos: Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend co…

@erikvoorhees

GLM 5.2 available on Venice, completely private and encrypted Best open source model for agents today

@teortaxestex

> GLM-5.2 scores 1524 on GDPval-AA v2 In retrospect I should have known that GLM also works on GDPeval, so there was no reason to expect it to flop there. Yeah peope saying 51-52 were right. This is about fair. It is stronger than Gemini

@andrew_n_carr

The importance of the glm 5.2 release is not about a "local" model near the frontier. It's about a small team, using that model, being able to build a system that trains models near the frontier.

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive