Z.ai's GLM-5.2 Beats Opus 4.8, Humans on Backend Take-Home in Offmute-v2 Release
A new offmute-v2 release added another performance claim for Z.ai, the Chinese AI lab behind GLM-5.2, a model released with public weights. The release described GLM-5.2 as outperforming Anthropic's Claude Opus 4.8 and human participants on the team's backend take-home test, while follow-up posts said it did so with higher-quality output, fewer tokens and lower cost. The same release presented offmute-v2 as a step forward for multi-stage media-to-transcript work.
Unlike earlier third-party rankings for GLM-5.2, the backend take-home result was presented by the team behind the test. Independent results this week put GLM-5.2 first among open-weight models on Vals AI's coding and agent benchmarks, including 64% on Vibe Code Bench, and on Artificial Analysis's Intelligence Index with a score of 51, while Anthropic's Claude Opus 4.8 remained ahead overall on AA-Briefcase, a long-horizon knowledge-work benchmark, and KernelBench-Mega, a GPU kernel-coding test.
From the sources (25 posts)
@valsaiGLM 5.2 had major improvements from its predecessor, GLM 5.1, with a 13% gain on the Vals Index and a 31% jump on Vibe Code Bench It also took the #1 spot on Terminal Bench 2.1 from Kimi K2.7 Code, an 11% improvement over GLM 5.1 https://t
@valsaiGLM 5.2 is the new open-weight SOTA on the Vals Index, Vibe Code Bench and Terminal Bench! It is also #5 across all models, and right on the heels of Opus 4.7 - released only two months ago
@artificialanlysContext on the result: CritPt is hard. It focuses on frontier physics problems developed by Argonne and UIUC through contributions from 60+ researchers globally, with the answer key and grading kept private. Models are independently benchma
@artificialanlysA standout number in Z ai’s GLM-5.2 launch is CritPt, a benchmark of unpublished research-level physics problems where it ties with Claude Opus 4.8 and is well above other open weights models Key takeaways: ➤ @Zai_org ’s GLM-5.2 (max reas
@zai_orgRT @opencode: GLM-5.2 now available in Go text · 1M context · same pricing as 5.1
@mtsliveSITUATION EXPLAINED: Z AI, a Chinese lab, just released a near-frontier model. • Competitive with everything except the very latest Opus, Fable, and GPT models • Pricing: $1 to $4 per million tokens, Haiku-tier • First open source Chinese
@zephyr_z9Massive I can't believe that a 700B model can hit this
@hosseebRT @Shaughnessy119: I agree with @hosseeb, the gap between Open and Closed source is not insane GLM 5.2 is 90% cheaper vs Fable 5, was rel…
@wesrothGLM-5.2 reached first place on Design Arena with an Elo score of 1360. The model gained 27 Elo points and climbed four positions, moving ahead of the now-unavailable Claude Fable 5.
@mtsliveSITUATION ANALYSIS: New Model From Z AI? Today, Chinese startup Z AI just released GLM-5.2, an open-source model closer to the frontier than any other in history. I know you’re probably used to hearing about benchmaxxed open-source slop mo
@chamathIt may be overfit so it’s a bit early to declare any sort of victory but it is really good to see open-weight models catch up to the closed source labs so quickly after their latest releases. With recursive RL, the delta in time between a
@blackhcRT @ProximalHQ: GLM 5.2 ranks #3 on FrontierSWE. It is only behind Fable 5 and Opus 4.8, and it outperforms GPT-5.5. This is the first mod…
@tftc21China just shipped a 744B parameter frontier AI model. MIT license, runs on consumer hardware, no NVIDIA required, 1M token context window. Dropped one day after the US suspended access to Fable 5. Open-source AI is freedom tech on th
@kimmonismusGLM-5.2 max is currently the third best model available, across both open and proprietary options. And that's fantastic. Open source is fundamentally important and must continue to hold a strong position so that everyone has open alternati
@andykonwinskiRT @ml_angelopoulos: Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend co…
@erikvoorheesGLM 5.2 available on Venice, completely private and encrypted Best open source model for agents today
@teortaxestex> GLM-5.2 scores 1524 on GDPval-AA v2 In retrospect I should have known that GLM also works on GDPeval, so there was no reason to expect it to flop there. Yeah peope saying 51-52 were right. This is about fair. It is stronger than Gemini
@andrew_n_carrThe importance of the glm 5.2 release is not about a "local" model near the frontier. It's about a small team, using that model, being able to build a system that trains models near the frontier.
@hosseebFor a long time I've been saying that the gap between open source and closed models is going to widen because of the data gap, hardware gap, and increased restrictions on distillation. I was wrong. is on another l
@gmoneynft@BitGrateful glm 5.2 is pretty close
@artificialanlys@jeremyphoward @Zai_org GLM-5.2 is between GPT-5.5 and Opus 4.8 in our new agentic knowledge work eval that we released just today. Very impressive
@nandodfRT @rasbt: Just caught up with the recent GLM-5.2 release. The best open-weight model today. Architecture-wise, it's build on the GLM-5 an…
@_arohan_RT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…
@andrewcurran_Rave reviews all over my timeline for GLM 5.2 the last 48 hours, open source is having an exceptional month.
@clementdelangue“Cost per task varies by ~800x across models tested: Claude Fable 5 leads the benchmark but costs more than $31 per task on average, compared to ~$0.04 for DeepSeek V4 Flash (max). The strongest price/performance options are open weights mo