Release Date: Jul 9, 2025

Developer SpaceXAIย ๐Ÿ‡บ๐Ÿ‡ธ
Context Window 256k
Max Output Tokens 128k
Token Costs (in/out) $3.00/15.00
Weights Private
Input Modalities

Accuracy

59.93 %

Avg. Cost (In/Out)

$ 3.00 / $ 15.00

Latency

8 min 51 s

Vals Index
BenchmarksAccuracyRankings

0.0%

ยฑ2.21
55/86

0.0%

ยฑ2.08
47/87

0.0%

ยฑ0.96
86/96

0.0%

ยฑ2.98
73/77

0.0%

ยฑ0.96
114/141

0.0%

ยฑ1.63
33/135

0.0%

ยฑ6.76
18/62

0.0%

ยฑ1.03
48/140

0.0%

ยฑ0.47
48/139

0.0%

ยฑ0.35
56/135

0.0%

ยฑ1.02
51/90

0.0%

ยฑ2.21
77/86
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : SpaceXAI
Temperature: 0.7
Top P: 0.95
Top K: Default
Max Output Tokens: 128,000

Updates

Jul 17, 2025

We evaluated Grok 4 on the Finance Agent, CorpFin, SWE-bench Verified, and LegalBench benchmarks and found strong results, especially on our private benchmarks.

Jul 13, 2025

In the livestream, Elon Musk called Grok 4 โ€œpartially blindโ€. We tested this claim on our two multimodal benchmarks (Mortgage Tax and MMMU Pro) and found a bigger gap between public and private benchmarks. We found that Grok 4 struggles to recognize unseen images, highlighting the importance of high-quality private datasets to evaluate image recognition capabilities.

As we continue to evaluate Grok 4 on our benchmarks, the model continues to struggle on our private ones. The middling performance on Tax Eval (67.6%) and Mortgage Tax (57.5%) is consistent with previous findings on our private legal tasks like Case Law and Contract Law.

On public benchmarks, Grok 4 achieves top-10 performance on both MMLU Pro (85.3%) and MMMU Pro (76.5%).

Jul 11, 2025

We found that Grok 4 struggles on our private benchmarks, in contrast to SOTA performance on AIME, Math 500, and GPQA Diamond.

Grok 4 delivers middle-of-the-pack performance on our private legal benchmarks. The model scores 80.6% on Case Law and 66.0% on Contract Law, underperforming Grok 3 Mini Reasoning on both and Grok 3 on Case Law. Notably, Grok 3 remains our top performer on the Case Law benchmark.

On public benchmarks, Grok 4 barely cracks the top 10 on MedQA at 92.5%, narrowly outperforming Grok 2. On MGSM, it fails to break the top 10 with 90.9%. This contrasts its SOTA performance on Math 500, suggesting Grok 4 struggles more with language than mathematical reasoning.

Jul 9, 2025

We received early access to SpaceXAIโ€™s latest Grok 4 and an initial set of smaller benchmarks. These early results show incredible performance โ€” the model sets the new state-of-the-art on AIME, GPQA Diamond, and Math 500 benchmarks! Grok 4 is extremely capable in its ability to answer challenging math and science questions.

We are continuing to run our evaluations on our private benchmarks and will release results shortly.