Dec 23, 2025
Full Results Released for GLM 4.7
Evaluations are finished for GLM 4.7.
- It is neck and neck with DeepSeek V3.2 (Thinking) on SWE-bench Verified, 67.0% and 68.8% respectively. This is a +11% bump on SWE-bench Verified compared to GLM 4.6
- On Terminal-Bench, it was one of the top open-weight models, along with DeepSeek V3.2 (Nonthinking)
- Overall, it is more token efficient than its predecessor, and was much cheaper per-task, despite having the same input and output pricing.
The model was tested with temperature=1 and default top_p for all benchmarks but SWE-bench Verified and Terminal-Bench, which used temperature=0.7 and top_p=1. Reasoning was enabled for all benchmarks.