Sep 21, 2026
xAI's Grok 4.7 evaluated across our benchmark suite
We evaluated xAIβs Grok 4.7 across 20 benchmarks plus the Vals Index.
-
#24 of 59 on the Vals Index (54.15%) at $4.78 per test, behind Grok 4.6 (59.17%) but ahead of Grok 4.5 (51.53%).
-
Strongest results on agentic legal and public-sector work: #5 of 63 on Harveyβs Legal Agent Benchmark (19.58%) and #5 of 37 on Public Benefits Bench (68.54%), plus #12 of 96 on MedScribe (87.21%) and #14 of 67 on Terminal-Bench 2.1 (76.03%).
-
Mid-board on Vibe Code Bench (75.86%, #20 of 97), Legal Research Bench (39.90%, #20 of 62), MedCode (48.69%, #22 of 94) and Finance Agent v2 (49.23%, #31 of 62). Weaker on EMB (54.94%, #35 of 59), IOI (39.39%, #22 of 29), ProofBench v1.1 (34.00%, #22 of 35) and SAGE (30.98%, #69 of 82).
-
Cost per test sits well below the frontier models on the Vals Index ($4.78 versus $18.81 to $28.92 for the top three), but climbs on long agentic tasks: $18.86 per test on ProgramBench (0.00%, one of 33 models yet to solve a task), $15.03 on Code Migration, $12.98 on Terminal-Bench Science (2.86%, #17 of 27) and $12.73 on Terminal-Bench 4 (12.12%, #16 of 29).
Grok 4.7 is priced at $2.00 per million input tokens and $6.00 per million output tokens. We evaluated it with reasoning effort set to xhigh.
Congrats to the xAI team on the release!