Command Palette
Search for a command to run...

Z.ai's GLM-5.2 Reaches Together AI and OpenRouter, Tops DeepSWE Open-Source Leaderboard at 44%

aiai-modelingai-open-modelsai-infrastructureai-inference-platformsai-research-evals 84 posts · 42 accounts

Chinese AI lab Z.ai's GLM-5.2, an AI model released with public model weights, is moving from benchmark leaderboards into wider deployment on hosted services and local setups. Together AI is serving the model on OpenRouter, a free US-hosted chat app is running on Together's platform, and a 467GB build has been uploaded for local use. DeepSWE, a software-engineering benchmark, put GLM-5.2 at the top of its open-source leaderboard with a 44% pass rate at max effort, 17 percentage points ahead of Kimi K2.7 Code.

The rollout extends GLM-5.2's reach beyond the benchmark streak that had already made it the highest-ranked open-weight model on Artificial Analysis's Intelligence Index and the only one above 60% on Vals AI's Vibe Code Bench, a test of whether models can build web applications from scratch. Code Arena's frontend leaderboard later placed GLM-5.2 Max second, 29 points ahead of Claude Opus 4.7 (Thinking) and behind only Anthropic's Claude Fable 5.

From the sources (25 posts)

@artificialanlys

@jeremyphoward @Zai_org GLM-5.2 is between GPT-5.5 and Opus 4.8 in our new agentic knowledge work eval that we released just today. Very impressive

@nandodf

RT @rasbt: Just caught up with the recent GLM-5.2 release. The best open-weight model today. Architecture-wise, it's build on the GLM-5 an…

@_arohan_

RT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…

@andrewcurran_

Rave reviews all over my timeline for GLM 5.2 the last 48 hours, open source is having an exceptional month.

@clementdelangue

“Cost per task varies by ~800x across models tested: Claude Fable 5 leads the benchmark but costs more than $31 per task on average, compared to ~$0.04 for DeepSeek V4 Flash (max). The strongest price/performance options are open weights mo

@chris_j_paxton

RT @jietang: GLM-5.2 is Fully Open, Frontier Intelligence Belongs to Everyone Today, the sudden restriction of certain frontier models is…

@clementdelangue

RT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…

@ollama

RT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…

@clementdelangue

RT @AndrewCurran_: Rave reviews all over my timeline for GLM 5.2 the last 48 hours, open source is having an exceptional month.

@zai_org

RT @ArtificialAnlys: Announcing AA-Briefcase, the benchmark for the next era of agentic knowledge work AA-Briefcase is our new benchmark f…

@andersonbcdefg

RT @teortaxesTex: ok, here's how GLM 5.2 performs on a bench it definitely didn't see, and where GLM 5.1 scored 0.0%. Closer to Opus 4.8 th…

@matvelloso

All day using GLM 5.2. Didn't miss much. First open model that passes the bar as a daily driver. Things are not going to be the same. Damn, now I want to buy some serious hardware.

@wesroth

RT @WesRoth: GLM-5.2 Max reached second place in the Code Arena Frontend leaderboard, behind only the currently unavailable Claude Fable 5.…

@quixiai

RT @jeremyphoward: Wow. @Zai_org GLM 5.2 is a marvel! It is *at least* as good as Opus 4.8 and GPT 5.5. It's super fast, inexpensive, and…

@zai_org

RT @ZixuanLi_: GLM-5.2 delivers a substantial leap in app development capabilities, which also represent demanding long-horizon tasks. Res…

@zai_org

Long-horizon is more than a concept. It should live in real-world scenarios, empowering AI builders to solve the problems that matter. And more scenarios are on the way.

@teortaxestex

Zhipu post-training team is very very good

@zai_org

RT @CunxiangWang: GLM-5.2 is not only stronger on benchmarks, but also much better in real app development scenarios — iOS, Android, WeChat…

@teortaxestex

GLM results are insanely high If this is "distillation", it's doing better than any previous attempt Eg ProofBench should be amenable both to distillation and RLVR, yet V4, Kimi, Mimo, Grok lol are all hopeless. Zhipu will obviously aim to

@teortaxestex

RT @banteg: not a bad result from glm-5.2, it found this after gpt-5.5 xhigh

@quixiai

RT @AlexFinn: I can't believe this is real I have GLM 5.2 running 100% locally on my Mac Studio. 2 bit quant. The results I'm getting are…

@vipulved

Spent the last 3 hours playing with GLM-5.2. It’s truly a stunning model. What a time to be alive.

@ollama

RT @matvelloso: All day using GLM 5.2. Didn't miss much. First open model that passes the bar as a daily driver. Things are not going to be…

@matansf

RT @alphatozeta8148: @baseten model performance team is absolutely cracked. @Zai_org GLM 5.2 is now 4x faster running at full 1M context!…

@teortaxestex

Interesting that in "GLM 5.2 is on par with X model", X is so widely distributed over different capabilities, and it's not like the gap is just greater for the difficult ones. It plainly can't do some parlor tricks Sonnet 3.5 pulled off. In

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive