Release Date: Apr 29, 2026

Developer Mistralย ๐Ÿ‡ซ๐Ÿ‡ท
Context Window 262k
Max Output Tokens 80k
Token Costs (in/out) $1.50/7.50
Weights Open
Input Modalities

Accuracy

17.95 % ยฑ 0.69

Cost / Test (Vals Index)

$ 8.010

Latency

57 min 26 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ0.69
44/46

0.0%

ยฑ1.37
45/49

0.0%

ยฑ1.64
45/46

0.0%

ยฑ0.97
42/49

0.0%

ยฑ2.00
44/49

0.0%

ยฑ1.15
67/85

0.0%

ยฑ2.01
78/84

0.0%

ยฑ0.84
92/95

0.0%

ยฑ2.86
21/23

0.0%

ยฑ3.36
51/76

0.0%

ยฑ0.92
100/140

0.0%

ยฑ1.71
79/84

0.0%

ยฑ2.39
129/133

0.0%

ยฑ0.42
106/133

0.0%

ยฑ2.11
68/83

0.0%

ยฑ2.70
51/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Mistral AI
Temperature: 0.7
Top P: 0.95
Top K: Default
Max Output Tokens: 80,000
Reasoning Effort: high

Updates

May 1, 2026

We evaluated Mistral Medium 3.5 on our suite of benchmarks. Here are the key takeaways:

  • Mistralโ€™s new reasoning-mode Medium 3.5 lands at #32 of 46 on the Vals Index (52.77%), and #10 of 18 among open-weight models.
  • It is a sizable jump over the prior Mistral Large 3 across most of the suite: +28 points on Finance Agent (46.1% vs 18.1%), +29 points on the SWE-bench Verified subset of the Vals Index (64.7% vs 35.3%), +21 points on Terminal-Bench 2.0 (30.3% vs 9.0%), and +13 points on SAGE (37.6% vs 24.6%).
  • One regression worth flagging: Case Law (v2) drops 17 points (44.2% vs 61.4%).

All evals were run via the Mistral API at temperature 0.7, top-p 0.95, with reasoning_effort: high and an 80k max output budget.