Release Date: Jul 9, 2025

Developer SpaceXAIย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 256k
Max Output Tokens 128k
Token Costs (in/out) $3.00/15.00
Weights Private
Input Modalities

Accuracy

62.74 %

Avg. Cost (In/Out)

$ 3.00 / $ 15.00

Latency

3 min 40 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ2.21
60/92

0.0%

ยฑ2.08
53/94

0.0%

ยฑ0.96
88/98

0.0%

ยฑ2.98
77/81

0.0%

ยฑ0.96
118/145

0.0%

ยฑ1.63
35/138

0.0%

ยฑ1.03
50/143

0.0%

ยฑ0.47
52/144

0.0%

ยฑ0.35
59/138

0.0%

ยฑ1.02
54/93

0.0%

ยฑ2.21
79/88

0.0%

ยฑ4.79
49/67
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : SpaceXAI
Temperature: 0.7
Top P: 0.95
Top K: Default
Max Output Tokens: 128,000

Updates

Jul 17, 2025

We evaluated Grok 4 on the Finance Agent, CorpFin, SWE-bench Verified, and LegalBench benchmarks and found strong results, especially on our private benchmarks.

Jul 13, 2025

In the livestream, Elon Musk called Grok 4 โ€œpartially blindโ€. We tested this claim on our two multimodal benchmarks (Mortgage Tax and MMMU Pro) and found a bigger gap between public and private benchmarks. We found that Grok 4 struggles to recognize unseen images, highlighting the importance of high-quality private datasets to evaluate image recognition capabilities.

As we continue to evaluate Grok 4 on our benchmarks, the model continues to struggle on our private ones. The middling performance on Tax Eval (67.6%) and Mortgage Tax (57.5%) is consistent with previous findings on our private legal tasks like Case Law and Contract Law.

On public benchmarks, Grok 4 achieves top-10 performance on both MMLU Pro (85.3%) and MMMU Pro (76.5%).

Jul 11, 2025

We found that Grok 4 struggles on our private benchmarks, in contrast to SOTA performance on AIME, Math 500, and GPQA Diamond.

Grok 4 delivers middle-of-the-pack performance on our private legal benchmarks. The model scores 80.6% on Case Law and 66.0% on Contract Law, underperforming Grok 3 Mini Reasoning on both and Grok 3 on Case Law. Notably, Grok 3 remains our top performer on the Case Law benchmark.

On public benchmarks, Grok 4 barely cracks the top 10 on MedQA at 92.5%, narrowly outperforming Grok 2. On MGSM, it fails to break the top 10 with 90.9%. This contrasts its SOTA performance on Math 500, suggesting Grok 4 struggles more with language than mathematical reasoning.

Jul 9, 2025

We received early access to SpaceXAIโ€™s latest Grok 4 and an initial set of smaller benchmarks. These early results show incredible performance โ€” the model sets the new state-of-the-art on AIME, GPQA Diamond, and Math 500 benchmarks! Grok 4 is extremely capable in its ability to answer challenging math and science questions.

We are continuing to run our evaluations on our private benchmarks and will release results shortly.