May 22, 2026
Alibaba's Qwen 3.7 Max evaluated on the Vals Index
-
We evaluated Alibaba’s new Qwen 3.7 Max on the Vals Index, where it places #5 with a score of 57.29%.
-
On SWE-bench Verified, Qwen 3.7 Max scores 68.8% overall and 66.7% on our Vals Index subset.
-
On CorpFin v2, the model scores 65.4% on the Vals Index task. Finance Agent v2 results are mid-tier at 48.4%.
-
On coding benchmarks, Qwen 3.7 Max scores 59.2% on Terminal Bench 2.0, 52.9% on the Vibe Code Bench Vals Index subset, and 26.0% on ProofBench.
The model has a ~1M-token context window. Evaluations were run with temperature 0.7 and preserve_thinking enabled.