Grok 4.20 (Reasoning)

Release Date: Mar 9, 2026

Developer SpaceXAIΒ πŸ‡ΊπŸ‡Έ
Context Window 2M
Max Output Tokens 2M
Token Costs (in/out) $2.00/6.00
Weights Private
Input Modalities

Accuracy

17.55 % Β± 0.53

Cost / Test (Vals Index)

$ 0.569

Latency

5 min 50 s

Vals Index
BenchmarksAccuracyRankings

0.0%

Β±0.53
45/46

0.0%

Β±0.11
47/48

0.0%

Β±1.55
44/46

0.0%

Β±0.32
46/49

0.0%

Β±2.41
39/49

0.0%

Β±2.12
72/84

0.0%

Β±2.10
80/83

0.0%

Β±0.99
83/94

0.0%

Β±3.42
48/75

0.0%

Β±0.86
37/139

0.0%

Β±2.06
75/84

0.0%

Β±1.59
30/132

0.0%

Β±7.49
16/62

0.0%

Β±1.03
37/137

0.0%

Β±0.48
96/136

0.0%

Β±0.34
40/132

0.0%

Β±0.89
27/88

0.0%

Β±2.01
51/82

0.0%

Β±0.99
46/54
Proprietary BenchmarksAcademic BenchmarksIndustry Partners
Vals
Default Provider : SpaceXAI
Temperature: 0.7
Top P: 0.95
Top K: Default
Max Output Tokens: 2,000,000

Updates

Mar 13, 2026

We evaluated Grok 4.20 (Reasoning) across our full benchmark suite.

Overall, the model generally is an improvement over previous SpaceXAI models, with room for improvement in certain domains.

Evaluations were run with a temperature of 0.7 and a top_p of 0.95 via the xAI API. This model is still in beta, and we will update results as and when updates are released by SpaceXAI.