Jun 17, 2026
z.AI's GLM 5.2 evaluated across our benchmark suite
-
z.AI just released their latest open-weights GLM 5.2 reasoning model. Itβs the #1 open-weight model on the Vals Index (65.02%, #5 overall), reclaiming the top open-weight spot from MiniMax-M3 β a 12.5-point jump over its predecessor, GLM 5.1 (52.45%).
-
Itβs strongest on coding: #1 open-weight and #3 overall on SWE-bench Verified (82.80%), and #1 open-weight on both Terminal-Bench 2.1 (67.79%) and Vibe Code Bench (63.96%) β the latter a 32-point leap from GLM 5.1.
-
It also leads open-weight models on agentic tasks, ranking #1 open-weight on Finance Agent v2 (49.70%) and #1 open-weight on Code Migration (37.87%).
-
On the index it comes in at $2.08/test β pricier than MiniMax-M3 ($1.50) and GLM 5.1 ($0.86), but well below frontier closed models like Claude Fable 5 ($5.16).
Eval settings: temperature=1, top_p=0.95, up to 131k max output tokens, run via the native z.AI API. Context window: 1M tokens.
Congrats to the team at z.AI on the release!