Release Date: May 7, 2025

Developer MistralΒ πŸ‡«πŸ‡·
Context Window 131k
Max Output Tokens 33k
Token Costs (in/out) $0.40/2.00
Weights Private
Input Modalities

Accuracy

58.64 %

Avg. Cost (In/Out)

$ 0.40 / $ 2.00

Latency

8.79 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.76
93/98

0.0%

Β±0.89
94/145

0.0%

Β±1.15
122/143

0.0%

Β±0.47
135/144

0.0%

Β±0.42
113/138

0.0%

Β±1.16
81/93
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : Mistral
Temperature: 0.7
Top P: Default
Top K: Default
Max Output Tokens: 32,768

Updates

May 9, 2025

We just evaluated Mistral Medium 3 on all benchmarks!

  • Mistral Medium 3 demonstrates consistent performance across both public and proprietary benchmarks, scoring 68.7% overall accuracy with strong results on CaseLaw (84.9%, #6/59) and Math500 (87.0%, #17/42) given its size and price.

  • The model outperforms Llama 4 Maverick (63.3% accuracy) in most benchmarks, particularly excelling in MGSM (91.6% vs 92.5%) and MMLU Pro (74.4% vs 79.4%).

  • While impressive, Mistral Medium 3 still trails behind Qwen 3 235B (81.0% accuracy) on several academic benchmarks, particularly Math500 (87.0% vs 94.6%) and AIME (42.3% vs 84.0%).

  • For users seeking speed-performance balance, Mistral Medium 3 offers good latency (14.37s) compared to Qwen 3 235B (94.31s), making it suitable for applications requiring faster response times while maintaining strong reasoning capabilities.