Grok 4.7 vs GPT-6 Sol: Benchmark Comparison

Grok 4.7 has the higher score on 7 of 21 shared benchmarks; GPT-6 Sol leads on 14.

The largest observed score gap is 57.00 pts on ProofBench v1.1 , where GPT-6 Sol leads.

Reported ±1 standard-error ranges overlap on 6 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Grok 4.7 GPT-6 Sol Gap Reported uncertainty
Vals Index 54.95% ±1.07 57.54% ±1.01 2.59 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 20.20% 28.11% 7.91 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 12.50% ±2.66 1.67% ±0.82 10.83 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 47.12% ±3.47 28.85% ±3.15 18.27 pts Reported ±1 SE ranges do not overlap
EMB 66.99% ±3.04 71.53% ±2.35 4.54 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 52.25% ±0.35 49.05% ±0.58 3.20 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 65.60% ±3.24 53.05% ±3.29 12.55 pts Reported ±1 SE ranges do not overlap
MedCode 49.55% ±2.17 47.07% ±2.12 2.48 pts Reported ±1 SE ranges overlap
MedScribe 89.38% ±1.89 82.03% ±1.94 7.34 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 26.00% ±4.41 83.00% ±3.77 57.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 69.26% ±1.33 74.81% ±2.43 5.56 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 25.23% ±2.92 30.18% ±3.09 4.95 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 10.00% ±3.61 30.00% ±5.52 20.00 pts Reported ±1 SE ranges do not overlap
SAGE 40.79% ±3.35 44.79% ±3.36 4.00 pts Reported ±1 SE ranges overlap
Code Migration 44.82% ±4.21 57.20% ±4.21 12.38 pts Reported ±1 SE ranges do not overlap
IOI 57.72% ±1.99 82.61% ±9.24 24.89 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.50% ±0.50 2.00% ±0.99 1.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 28.79% ±2.31 44.44% ±3.54 15.66 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 86.17% ±2.18 87.82% ±2.53 1.65 pts Reported ±1 SE ranges overlap
CyberBench v1.1 69.46% ±5.67 77.98% ±5.11 8.51 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 65.63% ±1.24 56.63% ±1.29 9.00 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Grok 4.7 average GPT-6 Sol average
Legal 29.81% 15.26%
Finance 61.61% 57.88%
Healthcare 69.47% 64.55%
Math 26.00% 83.00%
Science 34.83% 45.00%
Education 40.79% 44.79%
Coding 43.60% 54.81%
Cyber 69.46% 77.98%
Social Mobility 65.63% 56.63%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Grok 4.7 cost GPT-6 Sol cost Grok 4.7 latency GPT-6 Sol latency
Vals Index $12.12 $7.58 35m31s 29m11s
Vals RSI Index $237.12 $665.11 90h00m 90h00m
Harvey's Legal Agent Benchmark $11.13 $3.36 42m05s 13m40s
Legal Research Bench $4.92 $4.90 21m14s 24m09s
EMB $6.48 $1.28 34m57s 10m22s
Finance Agent (v2) $2.72 $2.12 15m33s 9m00s
Tax Agent Bench $1.80 $2.28 16m36s 17m32s
MedCode N/A N/A 3m33s 62.71s
MedScribe N/A N/A 2m33s 81.14s
ProofBench v1.1 $0.79 $0.48 12m47s 5m05s
BioMysteryBench $2.49 $0.66 11m37s 5m24s
MysteryMechanism $1.84 $0.46 17m33s 6m13s
Terminal-Bench Science $14.45 $5.82 1h29m 1h22m
SAGE N/A N/A 4m51s 31.42s
Code Migration $36.55 $15.68 1h09m 1h26m
IOI $12.71 $2.70 46m48s 30m38s
ProgramBench $46.49 $11.67 3h38m 46m11s
Terminal-Bench 4.0 $18.09 $5.79 50m21s 34m56s
Vibe Code Bench v1.1 $15.83 $26.36 35m47s 38m56s
CyberBench v1.1 $8.54 $1.34 15m30s 12m14s
Public Benefits Bench v1.1 $2.25 $2.49 16m57s 31m02s

Results available only for Grok 4.7

  • LegalBench

Results available only for GPT-6 Sol

None.

Model details Grok 4.7 Model details GPT-6 Sol