Artificial Analysis Ranks Z.ai GLM 5.2 as the Top Open Weights AI Model
Artificial Analysis, an AI analytics firm, ranked Z.ai's GLM 5.2 as the top open-weights language model as of June 30, 2026. The evaluation found the model scores 51 out of 100 on the firm's Intelligence Index, making it the strongest open-access option tracked. However, running the benchmark suite required approximately 141 million output tokens, with 95% dedicated to internal reasoning. That output volume significantly outpaces leading closed alternatives, generating nearly double GPT 5.5's 72 million tokens and exceeding Claude Opus 4.8's 117 million tokens.
The results highlight a trade-off between cognitive capability and computational efficiency at the frontier of generative AI. GLM 5.2 spent roughly 88 million tokens on a single reasoning test called Humanity's Last Exam and scored lower on a memory benchmark than both Opus 4.8 and GPT 5.5. Despite the verbosity, the model demonstrated competitiveness in specialized tasks. It matched Claude Opus 4.8 at a 21% score on CritPt, a physics reasoning benchmark developed by Argonne National Laboratory and the University of Illinois Chicago, and performed well on agentic workflow evaluations. A separate industry test by Baseten and Braintrust found the model to be 4.3 times cheaper for long-context retrieval tasks, with only a 3.5% drop in quality compared to Opus 4.8.
From the sources (5 posts)
@artificialanlysGLM-5.2 is the most intelligent open weights model available, but also the most verbose among the leading models GLM-5.2 (max) used ~141M output tokens (95% reasoning) to run the Artificial Analysis Intelligence Index (1.8x the average mod
@artificialanlysFrom GLM-5.1 to GLM-5.2, output to run the Artificial Analysis Intelligence Index rose from ~122M to ~141M, and the reasoning share climbed from 92% to 95%, while intelligence jumped 11 points (40 to 51). The newer version is both more cap
@artificialanlysThe thinking does pay off where reasoning is important. CritPt is a frontier physics benchmark developed by Argonne and UIUC, with contributions from 60+ researchers globally and applied by Artificial Analysis. On CritPt, GLM-5.2 ties Clau
@teortaxestexThe reason Anthropic strikes fear into the hearts of OpenAI TS is precisely the suspicion that no, GLM 5.2 10T would not be better than Fable 5, and neither would GPT 5.5 10T scaling laws optimized for *big* models I suspect "Fable" is not
@basetenWe partnered with Braintrust to highlight the real-world strength of the latest open-source model. For long-context retrieval tasks, GLM 5.2 is 4.3x cheaper with a mere 3.5% quality tradeoff vs. Opus 4.8 on Braintrust’s latest eval. Check