Archived Benchmark

Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases. Previous results are preserved here for posterity.

Academic

MGSM

Updated 1/9/2026

A multilingual benchmark for mathematical questions.

As of January 9, 2026, Claude Opus 4.5 (Thinking) ranks first on MGSM with 95.20%, followed by Claude Opus 4.5 (Nonthinking) at 94.76% and Claude Opus 4.1 (Thinking) at 94.44%.

MGSMMultilingual math word problems
ACCURACY

MGSM leaderboard

Rank Model Accuracy Cost In / Out Latency
1 Claude Opus 4.5 (Thinking) 95.20% $5 / $25 11.56s
2 Claude Opus 4.5 (Nonthinking) 94.76% $5 / $25 5.56s
3 Claude Opus 4.1 (Thinking) 94.44% $15 / $75 14.28s
4 Claude Sonnet 4.5 (Thinking) 94.33% $3 / $15 10.27s
5 Claude Opus 4.1 (Nonthinking) 94.22% $15 / $75 7.90s
6 GPT 5.2 94.00% $1.75 / $14 10.60s
7 Gemini 3 Pro (11/25) 93.93% $2 / $12 12.90s
8 Claude Opus 4 (Nonthinking) 93.78% $15 / $75 10.10s
9 o4 Mini 93.42% $1.1 / $4.4 8.07s
10 Gemini 3 Flash (12/25) 93.31% $0.5 / $3 14.02s
11 Claude Sonnet 4 (Nonthinking) 93.02% $3 / $15 5.99s
12 GPT 5.1 92.98% $1.25 / $10 8.35s
13 Claude 3.7 Sonnet (Thinking) 92.98% $3 / $15 23.51s
14 GPT 5 92.84% $1.25 / $10 20.35s
15 Claude 3.5 Sonnet Latest 92.58% $3 / $15 4.10s
16 GPT 5 Mini 92.58% $0.25 / $2 12.57s
17 Qwen 3 (235B) 92.47% $0.22 / $0.88 31.90s
18 Llama 4 Maverick 92.44% $0.22 / $0.88 2.71s
19 Claude 3.7 Sonnet (Nonthinking) 92.40% $3 / $15 4.86s
20 DeepSeek R1 92.25% $1.35 / $5.4 10.17s
21 Claude Haiku 4.5 (Thinking) 92.15% $1 / $5 9.62s
22 DeepSeek V3 92.15% $0.9 / $0.9 15.68s
23 Qwen 3 Max Preview 92.15% $1.2 / $6 21.16s
24 GPT OSS 120B 92.04% $0.15 / $0.6 5.79s
25 Qwen 3 Max 91.82% $1.2 / $6 11.45s
26 o3 91.75% $2 / $8 6.76s
27 DeepSeek V3 (03/24/2025) 91.67% $0.9 / $0.9 18.03s
28 Grok 3 91.35% $3 / $15 5.83s
29 o3 Mini 91.35% $1.1 / $4.4 15.98s
30 Llama 3.3 Instruct Turbo (70B) 91.09% $0.88 / $0.88 3.01s
31 Kimi K2 Instruct 90.95% $1 / $3 11.83s
32 Grok 4 90.91% $3 / $15 116.62s
33 Grok 4 Fast (Reasoning) 90.87% $0.2 / $0.5 3.85s
34 Mistral Medium 3.1 (05/2025) 90.87% $0.4 / $2 5.01s
35 DeepSeek V3.2 (Nonthinking) 90.87% $0.56 / $1.68 8.97s
36 Claude Sonnet 4 (Thinking) 90.87% $3 / $15 14.09s
37 GLM 4.5 90.84% $0.6 / $2.2 89.55s
38 GPT 4o (2024-08-06) 90.69% $2.5 / $10 6.73s
39 Grok 3 Mini Reasoning 90.44% $0.3 / $0.5 10.07s
40 GPT 4o (2024-11-20) 90.36% $2.5 / $10 3.98s
41 Grok 3 Mini Reasoning 90.36% $0.3 / $0.5 6.56s
42 Kimi K2 Thinking 90.15% $0.6 / $2.5 48.41s
43 Gemini 2.5 Flash Preview (9/25) (Nonthinking) 89.85% $0.3 / $2.5 3.41s
44 Gemini 2.5 Flash Preview (9/25) (Thinking) 89.82% $0.3 / $2.5 5.73s
45 GLM 4.6 89.75% $0.6 / $2.2 34.31s
46 Gemini 2.5 Flash Lite (9/25) (Nonthinking) 89.53% $0.1 / $0.4 1.47s
47 Grok 4.1 Fast (Reasoning) 89.53% $0.2 / $0.5 13.50s
48 o1 89.31% $15 / $60 11.21s
49 GPT 5 Nano 89.31% $0.05 / $0.4 22.77s
50 Gemini 1.5 Pro (002) 89.20% $1.25 / $5 2.80s
51 Gemini 2.0 Flash (001) 89.02% $0.1 / $0.4 1.54s
52 GPT OSS 20B 89.02% $0.07 / $0.3 5.79s
53 Gemini 2.5 Flash Lite (9/25) (Thinking) 88.40% $0.1 / $0.4 4.57s
54 GLM 4.7 88.18% $0.6 / $2.2 63.17s
55 Grok 4 Fast (Non-Reasoning) 88.00% $0.2 / $0.5 1.81s
56 Llama 4 Scout 87.96% $0.18 / $0.59 3.50s
57 MiniMax-M2.1 87.85% $0.3 / $1.2 13.86s
58 GPT 4.1 Mini 87.78% $0.4 / $1.6 2.75s
59 GPT 4.1 87.67% $2 / $8 2.19s
60 Grok 4.1 Fast Non-Reasoning 87.56% $0.2 / $0.5 3.68s
61 Mistral Large (11/2024) 87.24% $2 / $6 8.04s
62 Gemini 1.5 Flash (002) 86.58% $0.075 / $0.3 1.41s
63 Magistral Small 1.2 (09/2025) 86.25% $0.5 / $1.5 8.04s
64 GPT 4o Mini 86.18% $0.15 / $0.6 4.03s
65 Grok 2 86.15% $2 / $10 6.27s
66 DeepSeek V3.2 (Thinking) 86.04% $0.56 / $1.68 28.22s
67 Command A 85.71% $2.5 / $10 8.36s
68 Mistral Large 3 85.42% $0.5 / $1.5 7.62s
69 Claude 3.5 Haiku Latest 84.62% $0.8 / $4 3.74s
70 Mistral Small 3.1 (03/2025) 84.22% $0.075 / $0.3 3.87s
71 Mistral Small (02/2024) 83.96% $0.2 / $0.6 2.95s
72 Magistral Medium 1.2 (09/2025) 74.62% $2 / $5 11.93s
73 Jamba 1.6 Large 71.24% $2 / $8 9.86s
74 GPT 4.1 Nano 69.27% $0.1 / $0.4 1.46s
75 Jamba 1.6 Mini 41.71% $0.2 / $0.4 3.93s

Key Takeaways

  • In first place is Claude Opus 4.5 (Thinking) with 95.20% accuracy, but at a very high price point.
  • The majority of evaluated models achieve high accuracy, suggesting that this benchmark is reaching saturation. The model performances are very tightly clustered, with only marginal differences between them. There is likely not a statistically significant difference between the top models.
  • Across the board, all models exhibit better performance on the English version of the benchmark, underscoring the impact of pre-training data on performance.

Results

The evaluation of MGSM reveals the following:

  • Among the top models, performance is very tightly clustered, with only marginal differences between them.
  • Cheap models perform very comparatively to more expensive models, suggesting that premium models are not necessarily better for questions of this difficulty (or lack thereof).
  • Despite high overall performance, there is a noticeable performance drop in non-English languages. Specifically, models showed the lowest performance in Bengali, where the best result was 90.40% accuracy.

These results indicate that while state-of-the-art models excel in mathematical reasoning, language-specific performance discrepancies persist, likely due to the imbalance in training data across languages. For users seeking a model with strong multilingual mathematical capabilities, the benchmark provides a range of well-rounded options.

The high results also suggest that the models are potentially reaching saturation on this benchmark - and the benchmark may soon be unable to distinguish between models’ multilingual and mathematical reasoning capabilities.


Dataset and Context

The Multilingual Grade School Math Benchmark (MGSM) is an academic evaluation benchmark designed to assess the ability of language models to solve grade-school math problems in multiple languages. Derived from the well-known GSM8K dataset—which consists of 8.5K high-quality, diverse math word problems—the MGSM benchmark features a subset of 250 problems that have been carefully translated by human annotators into 10 typologically diverse languages (including underrepresented languages such as Bengali, Telugu, and Swahili). MGSM is commonly reported by model providers on new model releases.

Introduced in the paper Language Models are Multilingual Chain-of-Thought Reasoners by Shi et al. (2022) and supported by subsequent research on multilingual evaluation, the dataset not only measures a model’s numerical and reasoning capabilities but also its proficiency in processing linguistic variations across different scripts and cultural contexts.


Methodology

To ensure reproducible and fair comparisons, the evaluation of models on the MGSM benchmark was conducted using the same grading scripts provided in the OpenAI’s SimpleEvals GitHub repository. This approach guarantees consistency in the prompt format and testing environment across all languages.

We used the following prompt template to query each model:

Solve this math problem. Give the reasoning steps before giving the final answer on the last line by itself in the format of "Answer:". Do not add anything other than the integer answer after "Answer:".

{Question}

For non-English evaluations, the prompt was accurately translated into the target language while preserving the original structure. All models were tested with their default temperature settings.