Gemini 3.8 Flash vs Grok 4.7: Benchmark Comparison

Gemini 3.8 Flash has the higher score on 7 of 22 shared benchmarks; Grok 4.7 leads on 15.

The largest observed score gap is 25.71 pts on CyberBench v1.1 , where Grok 4.7 leads.

Reported ±1 standard-error ranges overlap on 11 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.8 Flash Grok 4.7 Gap Reported uncertainty
Vals Index 54.83% ±1.03 54.95% ±1.07 0.12 pts Reported ±1 SE ranges overlap
Vals RSI Index 19.93% 20.20% 0.27 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 10.00% ±2.42 12.50% ±2.66 2.50 pts Reported ±1 SE ranges overlap
Legal Research Bench 38.94% ±3.39 47.12% ±3.47 8.17 pts Reported ±1 SE ranges do not overlap
LegalBench 86.99% ±0.43 84.39% ±0.46 2.60 pts Reported ±1 SE ranges do not overlap
EMB 72.20% ±2.42 66.99% ±3.04 5.20 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 61.44% ±0.13 52.25% ±0.35 9.18 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 66.77% ±3.15 65.60% ±3.24 1.17 pts Reported ±1 SE ranges overlap
MedCode 48.13% ±2.18 49.55% ±2.17 1.42 pts Reported ±1 SE ranges overlap
MedScribe 84.50% ±1.94 89.38% ±1.89 4.88 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 48.00% ±5.02 26.00% ±4.41 22.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 62.22% ±1.70 69.26% ±1.33 7.04 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 36.49% ±3.24 25.23% ±2.92 11.26 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 8.57% ±3.37 10.00% ±3.61 1.43 pts Reported ±1 SE ranges overlap
SAGE 35.06% ±3.36 40.79% ±3.35 5.73 pts Reported ±1 SE ranges overlap
Code Migration 36.55% ±4.18 44.82% ±4.21 8.27 pts Reported ±1 SE ranges overlap
IOI 56.94% ±2.26 57.72% ±1.99 0.78 pts Reported ±1 SE ranges overlap
ProgramBench 1.00% ±0.70 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 19.19% ±2.52 28.79% ±2.31 9.60 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 78.65% ±3.88 86.17% ±2.18 7.52 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 43.75% ±2.21 69.46% ±5.67 25.71 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 65.29% ±1.24 65.63% ±1.24 0.34 pts Reported ±1 SE ranges overlap

Performance by category

Category Gemini 3.8 Flash average Grok 4.7 average
Legal 45.31% 48.00%
Finance 66.80% 61.61%
Healthcare 66.32% 69.47%
Math 48.00% 26.00%
Science 35.76% 34.83%
Education 35.06% 40.79%
Coding 38.47% 43.60%
Cyber 43.75% 69.46%
Social Mobility 65.29% 65.63%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.8 Flash cost Grok 4.7 cost Gemini 3.8 Flash latency Grok 4.7 latency
Vals Index $5.73 $12.12 51m57s 35m31s
Vals RSI Index $332.04 $237.12 90h00m 90h00m
Harvey's Legal Agent Benchmark $3.66 $11.13 29m21s 42m05s
Legal Research Bench $1.63 $4.92 5m24s 21m14s
LegalBench N/A N/A 3.32s 26.02s
EMB $8.24 $6.48 12m49s 34m57s
Finance Agent (v2) $2.00 $2.72 3m21s 15m33s
Tax Agent Bench $0.86 $1.80 2m58s 16m36s
MedCode N/A N/A 42.89s 3m33s
MedScribe N/A N/A 21.14s 2m33s
ProofBench v1.1 $0.60 $0.79 6m09s 12m47s
BioMysteryBench $1.55 $2.49 5m27s 11m37s
MysteryMechanism $1.73 $1.84 5m22s 17m33s
Terminal-Bench Science $5.64 $14.45 55m11s 1h29m
SAGE N/A N/A 28.64s 4m51s
Code Migration $18.49 $36.55 2h37m 1h09m
IOI $3.98 $12.71 14m20s 46m48s
ProgramBench $10.93 $46.49 46m39s 3h38m
Terminal-Bench 4.0 $8.77 $18.09 1h48m 50m21s
Vibe Code Bench v1.1 $6.87 $15.83 8m39s 35m47s
CyberBench v1.1 $1.31 $8.54 9m49s 15m30s
Public Benefits Bench v1.1 $1.01 $2.25 3m21s 16m57s

Results available only for Gemini 3.8 Flash

  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100
  • CUA-bench
  • Time Horizon Index: KSP

Results available only for Grok 4.7

None.

Model details Gemini 3.8 Flash Model details Grok 4.7