Jun 2, 2026
Alibaba's Qwen 3.7 Plus evaluated on the Vals Index
-
We evaluated Alibaba’s new Qwen 3.7 Plus on the Vals Index, where it ranks #13 with a score of 52.33%.
-
Its strongest Index component was Vibe Code Bench, where it ranks #8 on the Index subset with a score of 46.94%. On the full Vibe Code Bench leaderboard, the model scores 46.39%.
-
On coding benchmarks, Qwen 3.7 Plus matches Qwen 3.7 Max on the SWE-bench Verified Index subset at 66.67%, and scores 52.81% on Terminal Bench 2.1.
-
On finance benchmarks, it scores 62.01% on CorpFin v2 and 41.25% on Finance Agent v2, trailing Qwen 3.7 Max while slightly improving on Qwen 3.6 Plus in Finance Agent v2.
The model was run via the Alibaba API at temperature 0.7, with preserve_thinking enabled, a 1M-token context window, and 65,536 max output tokens.