Mar 13, 2026
Grok 4.20 Beta (03/09) evaluated on our full benchmark suite
We evaluated Grok 4.20 (Reasoning) across our full benchmark suite.
- Grok 4.20 (Reasoning) ranks #13/33 on Vals Index (57.70% accuracy) and #14/23 on Vals Multimodal Index (53.96%).
- It achieves strong results on academic and coding-heavy benchmarks, including #5 on AIME (96.46%), #6 on GPQA Diamond (88.64%), #9 on MMMU Pro (83.47%), and #9 on SWE-bench Verified (74.20%).
- On our finance and tax benchmarks, it delivers solid performance with #15 on Corp Fin (v2) (63.68%), #14 on Finance Agent (v1.1) (52.29%), and #16 on Tax Eval (v2) (74.12%).
- On healthcare tasks, performance is mixed: #16 on MedQA (94.55%), alongside #36 on MedCode (32.16%) and #43 on MedScribe (63.41%).
- On legal benchmarks, it currently ranks #30 on Case Law (v2) (54.45%) and #62 on LegalBench (77.74%).
Overall, the model generally is an improvement over previous SpaceXAI models, with room for improvement in certain domains.
Evaluations were run with a temperature of 0.7 and a top_p of 0.95 via the xAI API. This model is still in beta, and we will update results as and when updates are released by SpaceXAI.