Command Palette
Search for a command to run...

Z.ai's GLM-5.2 Beats Opus 4.8, Humans on Backend Take-Home in Offmute-v2 Release

aiai-modelingai-model-releasesai-open-modelsai-research-evals 78 posts · 37 accounts

A new offmute-v2 release added another performance claim for Z.ai, the Chinese AI lab behind GLM-5.2, a model released with public weights. The release described GLM-5.2 as outperforming Anthropic's Claude Opus 4.8 and human participants on the team's backend take-home test, while follow-up posts said it did so with higher-quality output, fewer tokens and lower cost. The same release presented offmute-v2 as a step forward for multi-stage media-to-transcript work.

Unlike earlier third-party rankings for GLM-5.2, the backend take-home result was presented by the team behind the test. Independent results this week put GLM-5.2 first among open-weight models on Vals AI's coding and agent benchmarks, including 64% on Vibe Code Bench, and on Artificial Analysis's Intelligence Index with a score of 51, while Anthropic's Claude Opus 4.8 remained ahead overall on AA-Briefcase, a long-horizon knowledge-work benchmark, and KernelBench-Mega, a GPU kernel-coding test.

From the sources (25 posts)

@valsai

GLM 5.2 had major improvements from its predecessor, GLM 5.1, with a 13% gain on the Vals Index and a 31% jump on Vibe Code Bench It also took the #1 spot on Terminal Bench 2.1 from Kimi K2.7 Code, an 11% improvement over GLM 5.1 https://t

@valsai

GLM 5.2 is the new open-weight SOTA on the Vals Index, Vibe Code Bench and Terminal Bench! It is also #5 across all models, and right on the heels of Opus 4.7 - released only two months ago

@artificialanlys

Context on the result: CritPt is hard. It focuses on frontier physics problems developed by Argonne and UIUC through contributions from 60+ researchers globally, with the answer key and grading kept private. Models are independently benchma

@artificialanlys

A standout number in Z ai’s GLM-5.2 launch is CritPt, a benchmark of unpublished research-level physics problems where it ties with Claude Opus 4.8 and is well above other open weights models Key takeaways: ➤ @Zai_org ’s GLM-5.2 (max reas

@zai_org

RT @opencode: GLM-5.2 now available in Go text · 1M context · same pricing as 5.1

@mtslive

SITUATION EXPLAINED: Z AI, a Chinese lab, just released a near-frontier model. • Competitive with everything except the very latest Opus, Fable, and GPT models • Pricing: $1 to $4 per million tokens, Haiku-tier • First open source Chinese

@zephyr_z9

Massive I can't believe that a 700B model can hit this

@hosseeb

RT @Shaughnessy119: I agree with @hosseeb, the gap between Open and Closed source is not insane GLM 5.2 is 90% cheaper vs Fable 5, was rel…

@wesroth

GLM-5.2 reached first place on Design Arena with an Elo score of 1360. The model gained 27 Elo points and climbed four positions, moving ahead of the now-unavailable Claude Fable 5.

@mtslive

SITUATION ANALYSIS: New Model From Z AI? Today, Chinese startup Z AI just released GLM-5.2, an open-source model closer to the frontier than any other in history. I know you’re probably used to hearing about benchmaxxed open-source slop mo

@chamath

It may be overfit so it’s a bit early to declare any sort of victory but it is really good to see open-weight models catch up to the closed source labs so quickly after their latest releases. With recursive RL, the delta in time between a

@blackhc

RT @ProximalHQ: GLM 5.2 ranks #3 on FrontierSWE. It is only behind Fable 5 and Opus 4.8, and it outperforms GPT-5.5. This is the first mod…

@tftc21

China just shipped a 744B parameter frontier AI model. MIT license, runs on consumer hardware, no NVIDIA required, 1M token context window. Dropped one day after the US suspended access to Fable 5. Open-source AI is freedom tech on th

@kimmonismus

GLM-5.2 max is currently the third best model available, across both open and proprietary options. And that's fantastic. Open source is fundamentally important and must continue to hold a strong position so that everyone has open alternati

@andykonwinski

RT @ml_angelopoulos: Just to be clear, if you remove Fable which is unavaialble, GLM-5.2 (Max) is the #1 model in the world for frontend co…

@erikvoorhees

GLM 5.2 available on Venice, completely private and encrypted Best open source model for agents today

@teortaxestex

> GLM-5.2 scores 1524 on GDPval-AA v2 In retrospect I should have known that GLM also works on GDPeval, so there was no reason to expect it to flop there. Yeah peope saying 51-52 were right. This is about fair. It is stronger than Gemini

@andrew_n_carr

The importance of the glm 5.2 release is not about a "local" model near the frontier. It's about a small team, using that model, being able to build a system that trains models near the frontier.

@hosseeb

For a long time I've been saying that the gap between open source and closed models is going to widen because of the data gap, hardware gap, and increased restrictions on distillation. I was wrong. is on another l

@gmoneynft

@BitGrateful glm 5.2 is pretty close

@artificialanlys

@jeremyphoward @Zai_org GLM-5.2 is between GPT-5.5 and Opus 4.8 in our new agentic knowledge work eval that we released just today. Very impressive

@nandodf

RT @rasbt: Just caught up with the recent GLM-5.2 release. The best open-weight model today. Architecture-wise, it's build on the GLM-5 an…

@_arohan_

RT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…

@andrewcurran_

Rave reviews all over my timeline for GLM 5.2 the last 48 hours, open source is having an exceptional month.

@clementdelangue

“Cost per task varies by ~800x across models tested: Claude Fable 5 leads the benchmark but costs more than $31 per task on average, compared to ~$0.04 for DeepSeek V4 Flash (max). The strongest price/performance options are open weights mo

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive