Archived Benchmark

Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases. Previous results are preserved here for posterity.

Academic

AIME

Updated 4/16/2026

Challenging national math exam given to top high-school students

As of April 16, 2026, Gemini 3.1 Pro Preview (02/26) ranks first on AIME with 98.13%, followed by GPT 5.2 (96.88%) and Muse Spark (96.88%).

AIMEElite high school math
ACCURACY

AIME leaderboard

Rank Model Accuracy Cost In / Out Latency
1 Gemini 3.1 Pro Preview (02/26) 98.13% $2 / $12 104.53s
2 GPT 5.2 96.88% $1.75 / $14 2m00s
3 Muse Spark 96.88% N/A 2m37s
4 Gemini 3 Pro (11/25) 96.68% $2 / $12 109.26s
5 GPT 5.4 (xhigh) 96.67% $2.5 / $15 2m01s
6 Grok 4.20 (Reasoning) 96.46% $2 / $6 106.80s
7 Claude Opus 4.7 96.25% $5 / $25 3m02s
8 Gemini 3 Flash (12/25) 95.63% $0.5 / $3 84.01s
9 Claude Opus 4.6 (Thinking) 95.63% $5 / $25 4m56s
10 GPT 5.4 Mini 95.63% $0.75 / $4.5 5m00s
11 Kimi K2.5 95.63% $0.6 / $3 5m23s
12 Claude Opus 4.5 (Thinking) 95.42% $5 / $25 3m32s
13 Qwen 3.6 Plus 94.58% $0.5 / $3 8m36s
14 GPT 5 93.37% $1.25 / $10 4m52s
15 GPT 5.1 93.33% $1.25 / $10 2m44s
16 GLM 4.7 93.33% $0.6 / $2.2 8m49s
17 GLM 4.6 92.71% $0.6 / $2.2 11m40s
18 GPT OSS 120B 92.60% $0.15 / $0.6 2m37s
19 Qwen 3.5 Flash 92.50% $0.1 / $0.4 4m00s
20 Claude Sonnet 4.6 92.29% $3 / $15 5m28s
21 Grok 4.1 Fast (Reasoning) 91.88% $0.2 / $0.5 92.00s
22 GLM 5.1 91.88% $1 / $3.2 6m15s
23 GLM 5 91.67% $1 / $3.2 7m39s
24 GPT 5 Mini 91.46% $0.25 / $2 105.53s
25 Grok 4 Fast (Reasoning) 91.25% $0.2 / $0.5 34.56s
26 MiniMax-M2.7 91.04% $0.3 / $1.2 3m57s
27 Grok 4 90.56% $3 / $15 2m13s
28 GPT 5.4 Nano 88.75% $0.2 / $1.25 65.88s
29 MiniMax-M2.5 88.75% $0.3 / $1.2 6m13s
30 Claude Sonnet 4.5 (Thinking) 88.19% $3 / $15 3m38s
31 GLM 4.5 86.67% $0.6 / $2.2 5m33s
32 o3 Mini 86.46% $1.1 / $4.4 2m35s
33 Qwen 3.5 Plus 86.04% $0.4 / $2.4 3m34s
34 GPT OSS 20B 86.04% $0.07 / $0.3 4m05s
35 Gemini 2.5 Pro Exp 85.83% $1.25 / $10 2m24s
36 Kimi K2 Thinking 85.42% $0.6 / $2.5 15m25s
37 o3 85.28% $2 / $8 4m26s
38 Grok 3 Mini Reasoning 85.00% $0.3 / $0.5 102.25s
39 DeepSeek V3.2 (Thinking) 84.58% $0.56 / $1.68 6m13s
40 Qwen 3 (235B) 83.96% $0.22 / $0.88 4m02s
41 o4 Mini 83.67% $1.1 / $4.4 55.11s
42 Magistral Medium 1.2 (09/2025) 83.54% $2 / $5 6m41s
43 Gemini 3.1 Flash Lite Preview 83.33% $0.25 / $1.5 40.56s
44 Claude Haiku 4.5 (Thinking) 82.71% $1 / $5 2m38s
45 GPT 5 Nano 81.18% $0.05 / $0.4 3m14s
46 Qwen 3 Max 81.04% $1.2 / $6 16.76s
47 Magistral Small 1.2 (09/2025) 80.68% $0.5 / $1.5 4m28s
48 Claude Opus 4.1 (Thinking) 78.18% $15 / $75 3m35s
49 MiniMax-M2.1 77.92% $0.3 / $1.2 3m58s
50 Claude Opus 4.5 (Nonthinking) 76.88% $5 / $25 17.44s
51 Claude Sonnet 4 (Thinking) 76.25% $3 / $15 4m32s
52 DeepSeek R1 73.96% $1.35 / $5.4 2m34s
53 o1 71.46% $15 / $60 2m57s
54 Grok 3 Mini Reasoning 70.63% $0.3 / $0.5 31.40s
55 DeepSeek V3.2 (Nonthinking) 64.79% $0.56 / $1.68 119.56s
56 Kimi K2 Instruct 62.71% $1 / $3 2m05s
57 Qwen 3 Max Preview 60.69% $1.2 / $6 3m27s
58 Grok 3 58.75% $3 / $15 63.99s
59 Llama 3.3 Nemotron Super (Thinking) 53.54% N/A 2m48s
60 DeepSeek V3 (03/24/2025) 52.20% $0.9 / $0.9 50.58s
61 Gemini 2.5 Flash Preview (9/25) (Thinking) 51.46% $0.3 / $2.5 61.60s
62 Gemini 2.5 Flash Preview (9/25) (Nonthinking) 49.79% $0.3 / $2.5 58.38s
63 GPT 4.1 Mini 49.38% $0.4 / $1.6 33.14s
64 Claude 3.7 Sonnet (Thinking) 44.58% $3 / $15 5m04s
65 Claude Opus 4.1 (Nonthinking) 44.24% $15 / $75 29.68s
66 Mistral Large 3 42.92% $0.5 / $1.5 104.33s
67 Mistral Medium 3.1 (05/2025) 42.29% $0.4 / $2 65.95s
68 Gemini 2.5 Flash Lite (9/25) (Thinking) 42.08% $0.1 / $0.4 51.15s
69 Claude Opus 4 (Nonthinking) 41.25% $15 / $75 37.03s
70 GPT 4.1 39.58% $2 / $8 2m41s
71 Claude Sonnet 4 (Nonthinking) 38.54% $3 / $15 23.40s
72 Grok 4 Fast (Non-Reasoning) 33.33% $0.2 / $0.5 5.88s
73 Gemini 2.0 Flash (001) 29.79% $0.1 / $0.4 11.21s
74 DeepSeek V3 27.50% $0.9 / $0.9 58.80s
75 Grok 4.1 Fast Non-Reasoning 27.08% $0.2 / $0.5 17.37s
76 GPT 4.1 Nano 26.46% $0.1 / $0.4 11.91s
77 Gemini 2.5 Flash Lite (9/25) (Nonthinking) 26.25% $0.1 / $0.4 18.70s
78 Llama 4 Maverick 25.21% $0.22 / $0.88 15.50s
79 Claude 3.7 Sonnet (Nonthinking) 22.29% $3 / $15 18.93s
80 Llama 4 Scout 18.96% $0.18 / $0.59 21.69s
81 Gemini 1.5 Pro (002) 18.75% $1.25 / $5 10.64s
82 Gemini 1.5 Flash (002) 17.29% $0.075 / $0.3 5.70s
83 Llama 3.3 Instruct Turbo (70B) 16.04% $0.88 / $0.88 11.24s
84 Grok 2 15.21% $2 / $10 57.88s
85 GPT 4o (2024-08-06) 13.96% $2.5 / $10 68.37s
86 Command A 13.33% $2.5 / $10 23.34s
87 GPT 4o (2024-11-20) 11.88% $2.5 / $10 15.78s
88 GPT 4o Mini 11.46% $0.15 / $0.6 28.77s
89 Claude 3.5 Sonnet Latest 10.00% $3 / $15 9.19s
90 Llama 3.3 Nemotron Super (Nonthinking) 9.38% N/A 15.85s
91 Mistral Large (11/2024) 9.17% $2 / $6 19.68s
92 Mistral Small (02/2024) 5.63% $0.2 / $0.6 13.23s
93 Mistral Small 3.1 (03/2025) 3.54% $0.075 / $0.3 11.68s
94 Claude 3.5 Haiku Latest 3.33% $0.8 / $4 9.05s
95 Jamba 1.6 Mini 0.42% $0.2 / $0.4 6.62s
96 Jamba 1.6 Large 0.42% $2 / $8 18.86s

Key Takeaways

  • Gemini 3.1 Pro Preview (02/26) is the new top-performing model on AIME at 98.12% accuracy.
  • Nine out of ten top models are reasoning models, underscoring the efficacy of reasoning for difficult math and coding problems.
  • Many of the questions are extremely resource intensive for top models - taking more than 30,000 reasoning tokens for a single question is very common, and generation can often take upwards of 15 minutes for the hardest questions.
  • As the AIME questions and answers are publicly available, there is a risk that models may have been exposed to them during pretraining. Notably, models tend to perform better on older (2024) questions compared to the newer 2025 set, raising questions about data contamination and true generalization.

Background

The American Invitational Mathematics Examination (AIME) is a prestigious, invite-only mathematics competition for high-school students who perform in the top 5% of the AMC 12 mathematics exam. It involves 15 questions of increasing difficulty, with the answer to every question being a single integer from 0 to 999. The median score is historically between 4 and 6 questions correct (out of the 15 possible). Two versions of the test are given every year (thirty questions total). You can view the questions from previous years on the AIME website

This examination serves as a crucial gateway for students aiming to qualify for the USA Mathematical Olympiad (USAMO). In general, the test is extremely challenging, and covers a wide range of mathematical topics, including algebra, geometry, and number theory.

The top models now achieve near-perfect accuracy on this benchmark, although performance can vary significantly between the 2024 and 2025 question sets.


Methodology

For this benchmark, we used the thirty questions from the 2024 and 2025 versions of the test (sixty questions total), modelling our approach after the repository from the GAIR NLP Lab.

To minimize parsing errors, we instructed the models with the following prompt template.

Please reason step by step, and put your final answer within \boxed{}

{Question}

The answer was then extracted from the boxed section and compared to the ground truth.

Although a few questions included an image or diagram, all of the information needed to solve the problem was present in the question text, so we did not include these images.

Reducing variance

Given the very low size of this benchmark, we ran each model 8 times on both AIME 2024 and AIME 2025 to reduce variance. We averaged the pass@1 performance across all runs for each model.