Release Date: Feb 16, 2026

Developer AlibabaΒ πŸ‡¨πŸ‡³
Context Window 991k
Max Output Tokens 66k
Token Costs (in/out) $0.40/2.40
Weights Private
Input Modalities

Accuracy

58.74 %

Avg. Cost (In/Out)

$ 0.40 / $ 2.40

Latency

10 min 36 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.97
65/98

0.0%

Β±3.15
69/80

0.0%

Β±3.18
69/93

0.0%

Β±1.67
38/138

0.0%

Β±1.00
34/143

0.0%

Β±0.41
23/142

0.0%

Β±0.33
34/138

0.0%

Β±1.01
93/93

0.0%

Β±2.03
59/88
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Alibaba
Temperature: 0.6
Top P: Default
Top K: Default
Max Output Tokens: 65,536

Updates

Feb 17, 2026

As we finish evaluating Qwen 3.5 Plus, here are our key takeaways:

  • The model performs well on knowledge benchmarks, ranking #6 on GPQA Diamond, #11 on MedQA and #9 on MMLU Pro, all of which are first place ranks among Open Weight models.
  • Qwen 3.5 Plus has a varied performance on our private benchmarks, ranking #30 on both the Mortgage Tax and SAGE benchmarks, but ranking #8 (and #1 among Open Weight models) on Finance Agent v1.1.
  • It loses out on performance due to having a highly sensitive content filter and having issues following a strict output format. This leads to low performance on the CaseLaw (v2) and MMMU Pro benchmarks.

Feb 16, 2026

We evaluated Qwen 3.5 Plus on the Vals Index. Here are the key takeaways:

  • Qwen 3.5 Plus places #10 overall on Vals Index (57.1% accuracy), and #3 among open-weight models.
  • The model is strongest on financial tasks: it places #6 on our Corp Fin (v2) and #8 on our Finance Agent subsets.
  • It is still weaker on legal and coding depth in this index mix, placing #25 on our Case Law (v2) and #17 on our SWE-bench Verified subsets.
  • It lands #11 on Terminal-Bench 2.0, which keeps it competitive on agentic coding tasks.