Archived Benchmark

Since performance on this benchmark has saturated, we no longer run this benchmark on new model releases. Previous results are preserved here for posterity.

Academic

MMLU Pro

Updated 9/1/2026

Academic multiple-choice benchmark covering 14 subjects including STEM, humanities, and social sciences.

As of September 1, 2026, Claude Fable 5.1 ranks first on MMLU Pro with 92.38%, followed by Claude Opus 5 (91.59%) and Claude Fable 5 (91.50%).

MMLU Pro14-subject academic knowledge test
ACCURACY

MMLU Pro leaderboard

Rank Model Accuracy Cost In / Out Latency
1 Claude Fable 5.1 92.38% $10 / $50 20.59s
2 Claude Opus 5 91.59% $5 / $25 9.60s
3 Claude Fable 5 91.50% $10 / $50 25.00s
4 Gemini 3.1 Pro Preview (02/26) 90.99% $2 / $12 23.91s
5 Gemini 3.8 Flash 90.22% $1.5 / $7.5 7.82s
6 Gemini 3.7 Flash 90.12% $1.5 / $7.5 4.16s
7 Gemini 3 Pro (11/25) 90.10% $2 / $12 24.27s
8 Claude Opus 4.7 89.87% $5 / $25 3m08s
9 Claude Opus 4.8 89.58% $5 / $25 23.44s
10 Gemini 3.5 Flash 89.52% $1.5 / $9 26.26s
11 Grok 4.6 89.40% $2 / $6 54.98s
12 Qwen 3.7 Max 89.31% $2.5 / $7.5 29.15s
13 Gemini 3.6 Flash 89.28% $1.5 / $7.5 7.30s
14 Grok 4.5 89.22% $2 / $6 5m58s
15 Claude Opus 4.6 (Thinking) 89.11% $5 / $25 49.45s
16 GPT-5.6 Sol 89.10% $4 / $20 15.65s
17 Muse Spark 1.1 88.73% $1.25 / $4.25 21.03s
18 Qwen 3.8 Max 88.60% $2 / $6 63.58s
19 Gemini 3 Flash (12/25) 88.59% $0.5 / $3 36.25s
20 Muse Spark 1.2 88.28% $1.25 / $4.25 44.23s
21 GPT 5.5 88.14% $5 / $30 42.11s
22 Kimi K3 87.97% $3 / $15 38.72s
23 Claude Opus 4.1 (Thinking) 87.92% $15 / $75 27.67s
24 Qwen 3.6 Plus 87.67% $0.5 / $3 46.18s
25 Kimi K2.6 87.57% $0.95 / $4 2m36s
26 Claude Sonnet 5 87.55% $2 / $10 25.19s
27 GPT 5.4 (xhigh) 87.48% $2.5 / $15 28.88s
28 Claude Sonnet 4.5 (Thinking) 87.36% $3 / $15 38.20s
29 Claude Sonnet 4.6 87.34% $3 / $15 55.41s
30 Muse Spark 87.32% N/A 46.24s
31 Claude Opus 4.5 (Thinking) 87.26% $5 / $25 38.92s
32 DeepSeek V4 87.25% $1.32 / $3.96 9m38s
33 Claude Opus 4.1 (Nonthinking) 87.21% $15 / $75 17.76s
34 Qwen 3.5 Plus 87.18% $0.4 / $2.4 76.68s
35 MiniMax-M2.1 87.05% $0.3 / $1.2 40.96s
36 DeepSeek V4 Pro 0813 86.97% $1.32 / $3.96 59.62s
37 GLM 5.1 86.90% $1 / $3.2 69.32s
38 GLM 5.3 86.77% $1.4 / $4.4 61.78s
39 GLM 5.2 86.71% $1.4 / $4.4 51.64s
40 GPT-5.6 Terra 86.66% $2 / $12 8.17s
41 GPT 5 86.54% $1.25 / $10 37.24s
42 GPT 5.1 86.38% $1.25 / $10 23.06s
43 Inkling 86.30% $1 / $4.05 73.26s
44 Grok 4.20 (Reasoning) 86.25% $2 / $6 11.65s
45 Gemini 3.1 Flash Lite Preview 86.24% $0.25 / $1.5 14.51s
46 GPT 5.2 86.23% $1.75 / $14 35.41s
47 DeepSeek V4 Flash 0731 86.21% $0.44 / $1.32 17.19s
48 Claude Opus 4 (Nonthinking) 86.17% $15 / $75 10.40s
49 GLM 5.3 Flash 86.06% $0.075 / $0.25 44.76s
50 GPT-5.6 Luna 86.04% $0.2 / $1.2 16.95s
51 GLM 5 86.03% $1 / $3.2 86.89s
52 Kimi K2.5 85.91% $0.6 / $3 39.20s
53 Grok 4.3 85.84% $1.25 / $2.5 10m37s
54 Gemini 3.5 Flash Lite 85.84% $0.3 / $2.5 4.82s
55 Nemotron 3 Ultra 85.76% N/A 32.92s
56 o3 85.59% $2 / $8 16.73s
57 Claude Opus 4.5 (Nonthinking) 85.59% $5 / $25 8.25s
58 Inkling Small 85.57% $0.3 / $1.2 71.69s
59 Grok 4 85.30% $3 / $15 89.25s
60 Qwen 3 Max Thinking 84.98% $1.2 / $6 110.21s
61 DeepSeek V3.2 (Thinking) 84.92% $0.56 / $1.68 90.15s
62 MiMo V2.5 Pro 84.59% $0.435 / $0.87 67.76s
63 GPT 5.4 Mini 84.55% $0.75 / $4.5 20.00s
64 Qwen 3 Max 84.36% $1.2 / $6 N/A
65 Qwen 3.8 27B 84.34% $0.5 / $3 69.85s
66 MiniMax-M3 84.22% $0.6 / $2.4 41.39s
67 Grok 4.1 Fast (Reasoning) 84.18% $0.2 / $0.5 25.47s
68 Qwen 3.5 Flash 84.06% $0.1 / $0.4 33.30s
69 Gemini 2.5 Pro Exp 84.06% $1.25 / $10 18.46s
70 Claude Sonnet 4 (Thinking) 83.86% $3 / $15 47.09s
71 Gemini 2.5 Flash Preview (9/25) (Nonthinking) 83.69% $0.3 / $2.5 8.53s
72 Gemini 2.5 Flash Preview (9/25) (Thinking) 83.66% $0.3 / $2.5 10.45s
73 Qwen 3 Max Preview 83.54% $1.2 / $6 43.43s
74 o1 83.49% $15 / $60 26.87s
75 DeepSeek R1 83.18% $1.35 / $5.4 27.28s
76 DeepSeek V3.2 (Nonthinking) 83.06% $0.56 / $1.68 40.77s
77 MiMo V2.5 82.93% $0.14 / $0.28 30.25s
78 GLM 4.7 82.74% $0.6 / $2.2 104.45s
79 Claude 3.7 Sonnet (Thinking) 82.73% $3 / $15 31.74s
80 GPT 5 Mini 82.23% $0.25 / $2 22.20s
81 GLM 4.6 82.20% $0.6 / $2.2 47.00s
82 Ling 3.0 Flash 82.01% $0.075 / $0.22 20.07s
83 Grok 3 Mini Reasoning 81.37% $0.3 / $0.5 16.89s
84 Qwen 3 (235B) 81.25% $0.22 / $0.88 52.35s
85 GLM 4.5 81.22% $0.6 / $2.2 2m17s
86 Kimi K2 Thinking 81.07% $0.6 / $2.5 94.77s
87 Claude 3.7 Sonnet (Nonthinking) 80.66% $3 / $15 6.30s
88 o4 Mini 80.56% $1.1 / $4.4 10.44s
89 GPT 4.1 80.50% $2 / $8 8.18s
90 MiniMax-M2.7 80.43% $0.3 / $1.2 50.01s
91 MiniMax-M2.5 80.09% $0.3 / $1.2 48.27s
92 Grok 3 Mini Reasoning 80.01% $0.3 / $0.5 7.43s
93 Grok 3 79.95% $3 / $15 12.10s
94 Mistral Large 3 79.82% $0.5 / $1.5 13.55s
95 Grok 4 Fast (Reasoning) 79.70% $0.2 / $0.5 7.34s
96 DeepSeek V3 (03/24/2025) 79.47% $0.9 / $0.9 25.05s
97 Claude Sonnet 4 (Nonthinking) 79.43% $3 / $15 9.71s
98 Llama 4 Maverick 79.42% $0.22 / $0.88 6.80s
99 Kimi K2 Instruct 79.39% $1 / $3 20.34s
100 GPT OSS 120B 79.17% $0.15 / $0.6 35.41s
101 Gemini 2.5 Flash Lite (9/25) (Thinking) 79.12% $0.1 / $0.4 9.62s
102 Claude Haiku 4.5 (Thinking) 78.72% $1 / $5 25.78s
103 o3 Mini 78.69% $1.1 / $4.4 22.38s
104 Gemini 2.5 Flash Lite (9/25) (Nonthinking) 78.64% $0.1 / $0.4 3.19s
105 Claude 3.5 Sonnet Latest 78.40% $3 / $15 6.29s
106 Gemini 2.0 Flash (001) 77.38% $0.1 / $0.4 4.32s
107 GPT 4.1 Mini 77.22% $0.4 / $1.6 3.83s
108 GPT 5.4 Nano 77.17% $0.2 / $1.25 8.21s
109 GPT 5 Nano 76.07% $0.05 / $0.4 26.11s
110 Grok 2 75.47% $2 / $10 8.75s
111 Mistral Medium 3.5 75.33% $1.5 / $7.5 66.38s
112 Gemini 1.5 Pro (002) 75.29% $1.25 / $5 3.34s
113 Mistral Medium 3.1 (05/2025) 75.29% $0.4 / $2 8.83s
114 Grok 4.1 Fast Non-Reasoning 75.21% $0.2 / $0.5 3.10s
115 GPT 4o (2024-08-06) 74.13% $2.5 / $10 9.00s
116 DeepSeek V3 73.82% $0.9 / $0.9 11.36s
117 GPT 4o (2024-11-20) 72.56% $2.5 / $10 9.14s
118 GPT OSS 20B 71.64% $0.07 / $0.3 30.21s
119 Llama 3.3 Nemotron Super (Nonthinking) 70.78% N/A 17.45s
120 Grok 4 Fast (Non-Reasoning) 70.34% $0.2 / $0.5 2.13s
121 Llama 3.3 Instruct Turbo (70B) 69.86% $0.88 / $0.88 4.18s
122 Mistral Large (11/2024) 69.71% $2 / $6 7.19s
123 Llama 4 Scout 69.63% $0.18 / $0.59 5.34s
124 Llama 3.3 Nemotron Super (Thinking) 69.58% N/A 36.35s
125 Command A 69.17% $2.5 / $10 9.85s
126 Laguna XS.2 69.05% N/A 47.99s
127 Laguna M.1 68.84% N/A 99.58s
128 Magistral Medium 1.2 (09/2025) 68.66% $2 / $5 31.99s
129 Mistral Small 3.1 (03/2025) 66.02% $0.075 / $0.3 3.60s
130 Gemini 1.5 Flash (002) 65.61% $0.075 / $0.3 1.68s
131 Mistral Small (02/2024) 64.44% $0.2 / $0.6 4.67s
132 Claude 3.5 Haiku Latest 64.12% $0.8 / $4 5.79s
133 GPT 4.1 Nano 63.48% $0.1 / $0.4 2.40s
134 GPT 4o Mini 62.73% $0.15 / $0.6 5.09s
135 Magistral Small 1.2 (09/2025) 62.13% $0.5 / $1.5 16.12s
136 Jamba 1.6 Large 49.78% $2 / $8 9.48s
137 Command R+ 44.00% $2.5 / $10 4.42s
138 Jamba 1.6 Mini 30.28% $0.2 / $0.4 2.43s

Key Takeaways

  • There is little left for providers to gain by hill-climbing MMLU-Pro: the top five models all sit above 90%, led by Claude Opus 5 at 91.59%.
  • The remaining weak spots are consistent and revealing — even the leader clears 95% on math and biology but drops to roughly 87% on law, history, and health, the more interpretive, judgment-heavy subjects.

Results


The results per subject are summarized in the graph below.


Dataset and Context

MMLU is a commonly-used academic evaluation that tests models on multiple-choice questions on subjects like physics, chemistry, etc. MMLU Pro is an improved version of MMLU, focusing on data quality and diversity. MMLU Pro consists of over 12,000 question-answer pairs, created through:

  • Filtering and enhancing classic MMLU questions
  • Expanding answer options from 4 to 10 choices
  • Incorporating new sources (STEM Website, TheoremQA, SciBench)
  • Expert verification of questions

The benchmark evaluates models on their ability to:

  • Demonstrate deep subject matter expertise
  • Apply complex reasoning to challenging problems
  • Show consistent performance across varied domains

Additional Notes

Methodology

All reported results use the 5-shot Chain-of-Thought prompting method, which included 5 examples per category in the prompt, as well as encouraging the models to think step by step. Few-shot CoT prompting is the approach used in the original paper. Here is the exact prompt:

The following are multiple-choice questions (with answers) about biology. Think step by step and then finish your answer with "The answer is (X)" where X is the correct letter choice.
Question: Which of the following represents an accurate statement concerning arthropods?
Options: A. They possess an exoskeleton composed primarily of peptidoglycan., B. They possess an open circulatory system with a dorsal heart., C. They are members of a biologically unsuccessful phylum incapable of exploiting diverse habitats and nutrition sources., D. They lack paired, jointed appendages.
Answer: Let's think step by step. Peptidoglycan is known to comprise the plasma membrane of most bacteria, rather than the exoskeleton of arthropods, which is made of chitin, which rules out (A). The answer (C) is false because arthropods are a highly successful phylum. Likewise, arthropods have paired, jointed appendages, which rules out (D). The only remaining option is (B), as arthropods have an open circulatory system with a dorsal tubular heart. The answer is (B).
Question: In a given population, 1 out of every 400 people has a cancer caused by a completely recessive allele, b. Assuming the population is in Hardy-Weinberg equilibrium, which of the following is the expected proportion of individuals who carry the b allele but are not expected to develop the cancer?
Options: A. 19/400, B. 1/400, C. 40/400, D. 38/400, E. 2/400, F. 1/200, G. 20/400, H. 50/400
Answer: Let's think step by step. According to the Hardy Weinberg Law, $p^2 + 2 p q + q^2 = 1$, and $p + q = 1$ where $p$ is the frequency of the dominant allele, $q$ is the frequency of the recessive allele, and $p^2$, $q^2$, and $2pq$ are the frequencies of dominant homozygous, recessive homozygous, and heterozygous individuals, respectively. ​The frequency of the recessive allele (q) is $\sqrt{\frac{1}{400}} = 0.05$. We have $p = 1 - q = 0.95$. The frequency of heterozygous individuals is $2pq = 2 \cdot 0.05 \cdot 0.95 = 0.095$. The number of heterozygous individuals is equal to the frequency of heterozygous individuals times the size of the population, or $0.095 * 400 = 38$. So we end up with 38/400. The answer is (D).
Question: A mutation in a bacterial enzyme changed a previously polar amino acid into a nonpolar amino acid. This amino acid was located at a site distant from the enzyme’s active site. How might this mutation alter the enzyme’s substrate specificity?
Options: A. By changing the enzyme’s pH optimum, B. By changing the enzyme's molecular weight, C. An amino acid change away from the active site increases the enzyme's substrate specificity., D. By changing the shape of the protein, E. By changing the enzyme's temperature optimum, F. By altering the enzyme's ability to be denatured, G. By changing the enzyme’s location in the cell, H. By changing the enzyme's color, I. An amino acid change away from the active site cannot alter the enzyme’s substrate specificity., J. By altering the enzyme's rate of reaction
Answer: Let's think step by step. A change in an amino acid leads to a change in the primary structure of the protein. A change in the primary structure may lead to a change in the secondary and the tertiary structure of the protein. A change in the tertiary structure means a change in the shape of the protein, so (C) has to be correct. Since the change does not affect the active site of the enzyme, we do not expect the activity of the enzyme to be affected. The answer is (D).
Question: Which of the following is not a way to form recombinant DNA?
Options: A. Translation, B. Conjugation, C. Specialized transduction, D. Transformation
Answer: Let's think step by step. The introduction of foreign DNA or RNA into bacteria or eukaryotic cells is a common technique in molecular biology and scientific research. There are multiple ways foreign DNA can be introduced into cells including transformation, transduction, conjugation, and transfection. In contrast, (A) is not a way to form DNA: during translation the ribosomes synthesize proteins from RNA. The answer is (A).
Question: Which of the following is not known to be involved in the control of cell division?
Options: A. Microtubules, B. Checkpoints, C. DNA polymerase, D. Centrosomes, E. Cyclins, F. Mitochondria, G. Protein kinases, H. Fibroblast cells
Answer: Let's think step by step. Normal cells move through the cell cycle in a regulated way. At the checkpoint stage, they use information about their own internal state and cues from the environment around them to decide whether to proceed with cell division. Cues like these act by changing the activity of core cell cycle regulators inside the cell. The most common regulators are cyclins and cyclin-dependent kinases. Fibroblast cells do not play any role in cell division. The answer is (H).

Which of the following would most likely provide examples of mitotic cell divisions?

A - cross section of muscle tissue
B - longitudinal section of a shoot tip
C - longitudinal section of a leaf vein
D - cross section of a fruit
E - cross section of a leaf
F - longitudinal section of a petal
G - longitudinal section of a seed
H - cross section of an anther (site of pollen production in a flower)

To grade the answers for correctness, we use the following regex, consistent with the original paper:

(?:answer is \(?(B)\)?)|(?:[Aa]nswer:\s*(B))

The overall score was calculated by averaging the 14 task-specific scores.

In total, the benchmark used ~15M input tokens per model.