Apr 4, 2025
Mistral Small 3.1 (2503) evaluated on all benchmarks!
We just evaluated Mistral Small 2503 on all benchmarks!
- Mistral Small 3.1 is Mistral AIβs latest small model, achieving an average accuracy of 61.4% across all benchmarks with a latency of 6.52s - faster than GPT-4o Mini (9.89s) and Llama 3.3 70B (7.67s).
- Despite its compact size, Mistral Small outperforms Claude 3.5 Haiku (60.2%) in overall accuracy while offering competitive performance to GPT-4o Mini (62.8%).
- The model excels on MGSM with 85.4% accuracy, comparable to Claude Haiku (85.9%) but behind Llama 3.3 70Bβs impressive 91.3%.
- Like Claude Haiku, the model struggles with AIME (both 3.5%), well behind GPT-4o Mini (11.5%) and Llama 3.3 70B (16.6%).