Z.ai's GLM-5.2 Scores 64% on Vibe Code Bench, Only Open-Weight Model Above 60%
Benchmark firm Vals AI ranked Chinese AI lab Z.ai's GLM-5.2, an AI model released with public weights, as the highest open-weight system across the Vals Index, Vibe Code Bench and Terminal Bench 2.1. On Vibe Code Bench v1.1, a test of whether models can build web applications from scratch, it scored 64%, making it the only open-weight model above 60% and 14 percentage points ahead of the next open-weight entry.
Separate KernelBench-Mega results, which test models writing GPU megakernels from scratch on Nvidia's RTX PRO 6000, H100 and B200 chips, put Claude Opus 4.8 from Anthropic first overall and GLM-5.2 first among open-weight models. Artificial Analysis had earlier ranked GLM-5.2 first among open-weight systems on its Intelligence Index with a score of 51, and developer tools including ncode and the Noumena platform have added support.
From the sources (25 posts)
@askalphaxivRT @askalphaxiv: Introducing GLM-5.2 for understanding research papers 🚀 Highlight any section of a paper to ask questions and “@” othe…
@theahmadosmanGLM 5.2
@victormustarRT @sam_paech: GLM-5.2 is an extremely strong open weights model. Kudos to the @Zai_org team.
@jeremyphowardRT @xeophon: 51 is the same score as GPT-5.4 xhigh, btw. The model was released 3 months ago and frontier at the time
@adinayakupReally cool to see the GLM 5.2 blog on @huggingface 🔥
@jeremyphowardRT @hallerite: GLM5.2 brings back the critic. It was just a matter of time until we people would realize that group-based variance reducti…
@jeremyphowardRT @ml_angelopoulos: Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend co…
@valsaiGLM 5.2 had major improvements from its predecessor, GLM 5.1, with a 13% gain on the Vals Index and a 31% jump on Vibe Code Bench It also took the #1 spot on Terminal Bench 2.1 from Kimi K2.7 Code, an 11% improvement over GLM 5.1 https://t
@valsaiGLM 5.2 is the new open-weight SOTA on the Vals Index, Vibe Code Bench and Terminal Bench! It is also #5 across all models, and right on the heels of Opus 4.7 - released only two months ago
@artificialanlysContext on the result: CritPt is hard. It focuses on frontier physics problems developed by Argonne and UIUC through contributions from 60+ researchers globally, with the answer key and grading kept private. Models are independently benchma
@artificialanlysA standout number in Z ai’s GLM-5.2 launch is CritPt, a benchmark of unpublished research-level physics problems where it ties with Claude Opus 4.8 and is well above other open weights models Key takeaways: ➤ @Zai_org ’s GLM-5.2 (max reas
@zai_orgRT @opencode: GLM-5.2 now available in Go text · 1M context · same pricing as 5.1
@mtsliveSITUATION EXPLAINED: Z AI, a Chinese lab, just released a near-frontier model. • Competitive with everything except the very latest Opus, Fable, and GPT models • Pricing: $1 to $4 per million tokens, Haiku-tier • First open source Chinese
@zephyr_z9Massive I can't believe that a 700B model can hit this
@hosseebRT @Shaughnessy119: I agree with @hosseeb, the gap between Open and Closed source is not insane GLM 5.2 is 90% cheaper vs Fable 5, was rel…
@wesrothGLM-5.2 reached first place on Design Arena with an Elo score of 1360. The model gained 27 Elo points and climbed four positions, moving ahead of the now-unavailable Claude Fable 5.
@mtsliveSITUATION ANALYSIS: New Model From Z AI? Today, Chinese startup Z AI just released GLM-5.2, an open-source model closer to the frontier than any other in history. I know you’re probably used to hearing about benchmaxxed open-source slop mo
@chamathIt may be overfit so it’s a bit early to declare any sort of victory but it is really good to see open-weight models catch up to the closed source labs so quickly after their latest releases. With recursive RL, the delta in time between a
@blackhcRT @ProximalHQ: GLM 5.2 ranks #3 on FrontierSWE. It is only behind Fable 5 and Opus 4.8, and it outperforms GPT-5.5. This is the first mod…
@tftc21China just shipped a 744B parameter frontier AI model. MIT license, runs on consumer hardware, no NVIDIA required, 1M token context window. Dropped one day after the US suspended access to Fable 5. Open-source AI is freedom tech on th
@kimmonismusGLM-5.2 max is currently the third best model available, across both open and proprietary options. And that's fantastic. Open source is fundamentally important and must continue to hold a strong position so that everyone has open alternati
@andykonwinskiRT @ml_angelopoulos: Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend co…
@erikvoorheesGLM 5.2 available on Venice, completely private and encrypted Best open source model for agents today
@teortaxestex> GLM-5.2 scores 1524 on GDPval-AA v2 In retrospect I should have known that GLM also works on GDPeval, so there was no reason to expect it to flop there. Yeah peope saying 51-52 were right. This is about fair. It is stronger than Gemini
@andrew_n_carrThe importance of the glm 5.2 release is not about a "local" model near the frontier. It's about a small team, using that model, being able to build a system that trains models near the frontier.