Archived Benchmark

Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases. Previous results are preserved here for posterity.

Academic

MATH 500

Updated 1/9/2026

Academic math benchmark on probability, algebra, and trigonometry

As of January 9, 2026, Gemini 3 Pro (11/25) ranks first on MATH 500 with 96.40%, followed by Grok 4 (96.20%) and GPT 5 (96.00%).

MATH 500Algebra, probability, trigonometry problems
ACCURACY

MATH 500 leaderboard

Rank Model Accuracy Cost In / Out Latency
1 Gemini 3 Pro (11/25) 96.40% $2 / $12 32.15s
2 Grok 4 96.20% $3 / $15 27.26s
3 GPT 5 96.00% $1.25 / $10 32.22s
4 Claude Opus 4.1 (Thinking) 95.40% $15 / $75 45.98s
5 Gemini 2.5 Pro Exp 95.20% $1.25 / $10 25.83s
6 GPT 5 Mini 94.80% $0.25 / $2 12.06s
7 GPT OSS 120B 94.80% $0.15 / $0.6 16.27s
8 o3 94.60% $2 / $8 16.59s
9 Qwen 3 (235B) 94.60% $0.22 / $0.88 2m23s
10 o4 Mini 94.20% $1.1 / $4.4 12.54s
11 Grok 3 Mini Reasoning 94.20% $0.3 / $0.5 22.77s
12 GPT OSS 20B 94.20% $0.07 / $0.3 31.47s
13 Kimi K2 Instruct 94.20% $1 / $3 37.35s
14 GLM 4.5 94.00% $0.6 / $2.2 84.93s
15 GPT 5 Nano 93.80% $0.05 / $0.4 22.40s
16 Claude Sonnet 4 (Thinking) 93.80% $3 / $15 63.47s
17 Claude Opus 4.1 (Nonthinking) 93.00% $15 / $75 30.93s
18 DeepSeek R1 92.20% $1.35 / $5.4 2m36s
19 o3 Mini 91.80% $1.1 / $4.4 14.37s
20 Gemini 2.5 Flash Preview 4/17 (Thinking) 91.80% $0.3 / $2.5 23.66s
21 Gemini 2.5 Flash Preview 4/17 (Nonthinking) 91.60% $0.3 / $2.5 9.50s
22 Claude 3.7 Sonnet (Thinking) 91.60% $3 / $15 94.24s
23 Llama 3.3 Nemotron Super (Thinking) 91.40% N/A 42.83s
24 Claude Opus 4 (Nonthinking) 90.40% $15 / $75 14.81s
25 o1 90.40% $15 / $60 23.54s
26 Claude Sonnet 4 (Nonthinking) 90.32% $3 / $15 10.03s
27 Grok 3 89.80% $3 / $15 9.09s
28 Gemini 2.0 Flash Exp 89.00% $0.075 / $0.3 3.35s
29 MiniMax-M2.1 89.00% $0.3 / $1.2 59.69s
30 DeepSeek V3 (03/24/2025) 88.60% $0.9 / $0.9 12.54s
31 Gemini 2.0 Flash (001) 88.00% $0.1 / $0.4 3.37s
32 GPT 4.1 Mini 88.00% $0.4 / $1.6 5.76s
33 GPT 4.1 87.20% $2 / $8 12.53s
34 Mistral Medium 3.1 (05/2025) 87.00% $0.4 / $2 14.13s
35 Llama 4 Maverick 85.20% $0.22 / $0.88 7.47s
36 Gemini 2.0 Flash Thinking Exp 84.60% $0.1 / $0.7 11.80s
37 Gemini 1.5 Pro (002) 82.80% $1.25 / $5 5.02s
38 DeepSeek V3 80.40% $0.9 / $0.9 7.94s
39 GPT 4.1 Nano 80.20% $0.1 / $0.4 3.37s
40 Llama 4 Scout 79.20% $0.18 / $0.59 10.52s
41 Gemini 1.5 Flash (002) 78.80% $0.075 / $0.3 2.65s
42 Grok 2 78.40% $2 / $10 20.44s
43 Claude 3.7 Sonnet (Nonthinking) 76.80% $3 / $15 5.53s
44 Command A 76.20% $2.5 / $10 8.66s
45 GPT 4o (2024-08-06) 75.20% $2.5 / $10 12.29s
46 Mistral Large (11/2024) 74.40% $2 / $6 9.93s
47 GPT 4o (2024-11-20) 74.00% $2.5 / $10 12.80s
48 Llama 3.3 Instruct Turbo (70B) 73.40% $0.88 / $0.88 5.41s
49 GPT 4o Mini 72.60% $0.15 / $0.6 6.30s
50 Claude 3.5 Sonnet Latest 72.40% $3 / $15 4.63s
51 Llama 3.1 Instruct Turbo (405B) 71.40% $3.5 / $3.5 45.30s
52 Llama 3.3 Nemotron Super (Nonthinking) 71.20% N/A 12.55s
53 Mistral Small (02/2024) 70.60% $0.2 / $0.6 4.88s
54 Grok 3 Mini Reasoning 70.20% $0.3 / $0.5 9.88s
55 Mistral Small 3.1 (03/2025) 68.40% $0.075 / $0.3 6.41s
56 Llama 3.1 Instruct Turbo (70B) 65.00% $0.88 / $0.88 9.20s
57 Claude 3.5 Haiku Latest 64.20% $0.8 / $4 5.13s
58 Jamba 1.6 Large 54.80% $2 / $8 13.01s
59 Llama 3.1 Instruct Turbo (8B) 44.40% $0.18 / $0.18 5.84s
60 Jamba 1.6 Mini 25.40% $0.2 / $0.4 4.86s

Key Takeaways

  • Gemini 3 Pro (11/25) is the best performing model, achieving a 96.40% accuracy. Grok 4 is a close second.
  • Most recent models now consistently score over 90%, indicating the strength of AI models in answering math questions. These questions are all in the public domain, creating a high risk of pre-training on the test set. It is possible their results no longer reflect a meaningful difference in model capabilities.

Results

In general, the models performed well on the MATH benchmark - the top models were able to solve over 90% of questions correctly. As this benchmark is public, and been out for several years, it does raise questions if the models have been pre-trained on the test data.

Some models performed similarly to others, but were much more verbose. For example, DeepSeek R1 used over 26,242 characters for one problem. In contrast, o3 Mini solved it in just just around 10,000 characters (2700 tokens).


Example

We provide examples of a question from the dataset, as well as the outputs from each model. You will see that each model chose different paths to solve the same problem. You will also see the logical deductions made by the models between each step in their attempts.

Question

Compute [1/cos^2(10°) + 1/sin^2(20°) + 1/sin^2(40°)]

We wish to simplify
S = 1⁄cos²10° + 1⁄sin²20° + 1⁄sin²40°.
Notice that cos 10° = sin 80° (since sin (90°–θ) = cos θ). Thus, we can rewrite the first term as

1⁄cos²10° = 1⁄sin²80°.

So the sum becomes

S = 1⁄sin²20° + 1⁄sin²40° + 1⁄sin²80°.

A useful identity is
csc²θ = 1 + cot²θ,
so we may write
S = [
1 + cot²20°
] + [
1 + cot²40°
] + [
1 + cot²80°
] = 3 + (cot²20° + cot²40° + cot²80°).

A numerical check (or a derivation using product–sum identities) shows that
cot²20° + cot²40° + cot²80° = 8.
(For example, one may verify using a calculator that cot 20° ≈ 2.7475 so its square is ≈7.547, cot 40° ≈ 1.1918 (square ≈1.420), and cot 80° ≈ 0.1763 (square ≈0.031) whose sum is approximately 8.)

Thus, we have
S = 3 + 8 = 11.

Therefore, the fully simplified solution is
\boxed{11}.

INCORRECT


Methodology

This benchmark is an adaptation of the MATH benchmark, first published in Measuring Mathematical Problem Solving With the MATH Dataset. The MATH benchmark is commonly reported on new model releases.

We sample 500 diverse problems from this benchmark - spanning topics like probability, algebra, trigonometry, and geometry. The questions are designed to test a model’s ability to apply mathematical principles, execute complex calculations, and communicate solutions clearly.

Unlike in original paper, which fine-tuned models to produce LaTeX output, we used the following prompt template to ensure the models produce outputs in the correct format.

Answer the following math question, given in LaTeX format, clearly and concisely, and present the final answer as \(\boxed{x}\), where X is the fully simplified solution.

Example:
**Question:** \(\int_0^1 (3x^2 + 2x) \,dx\)
**Solution:** \(\int (3x^2 + 2x) \,dx = x^3 + x^2 + C\) Evaluating from 0 to 1: \((1^3 + 1^2) - (0^3 + 0^2) = 1 + 1 - 0 = 2 \boxed{2}\)

Now, solve the following question: {question}

We also used the parsing logic from the the PRM800K dataset grader. We found that this was much more reliable in extracting and evaluating the model’s output - it was robust towards differing formats and mathematical formulations, compared to the parsing logic from the original MATH paper.

All models were evaluated with temperature set to 0, except for the reasoning models that force a certain temperature (like 3.7 at 1, or o1 not accepting temperature).