Oct 28, 2025
GLM 4.6 Evaluated on All Benchmarks!
We evaluated GLM 4.6 on all benchmarks and found it improves upon the already strong performance of GLM 4.5, leading the [open-source category of our Vals Index.
- While GLM 4.6 outperforms its predecessor GLM 4.5, it underperforms compared to Qwen 3 Max
- GLM 4.6 excels at coding, placing in the top 10 on both SWE-bench Verified, Terminal-Bench, and LiveCodeBench. This reproduces z.AI internal results on the latter two benchmarks, though lags 10 points behind internal numbers on SWE-bench Verified.
- GLM 4.6 struggles on our private benchmarks, underperforming compared to GLM 4.5 on our Finance Industry Leaderboard.
See our comparison with GLM 4.5 for more context.