Model comparison

Gemini 3.7 Flash vs Grok 4.7: Benchmark Comparison

Gemini 3.7 Flash has the higher score on 7 of 19 shared benchmarks; Grok 4.7 leads on 12.

The largest observed score gap is 32.00 pts on ProofBench v1.1 , where Gemini 3.7 Flash leads.

Reported ±1 standard-error ranges overlap on 6 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.7 Flash Grok 4.7 Gap Reported uncertainty
Vals Index 59.31% ±1.06 60.22% ±1.08 0.91 pts Reported ±1 SE ranges overlap
Vals RSI Index 18.27% 20.20% 1.93 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 8.75% ±2.15 12.50% ±2.42 3.75 pts Reported ±1 SE ranges overlap
Legal Research Bench 34.62% ±3.31 47.12% ±3.47 12.50 pts Reported ±1 SE ranges do not overlap
LegalBench 87.26% ±0.42 84.39% ±0.46 2.87 pts Reported ±1 SE ranges do not overlap
EMB 71.33% ±2.26 66.99% ±3.04 4.33 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 59.04% ±0.27 52.25% ±0.35 6.79 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 57.66% ±3.31 65.60% ±3.24 7.93 pts Reported ±1 SE ranges do not overlap
MedCode 53.39% ±2.12 49.55% ±2.17 3.84 pts Reported ±1 SE ranges overlap
MedScribe 83.94% ±2.00 89.38% ±1.89 5.44 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 58.00% ±4.96 26.00% ±4.41 32.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 5.71% ±2.79 11.43% ±3.83 5.71 pts Reported ±1 SE ranges overlap
SAGE 49.23% ±3.38 40.79% ±3.35 8.44 pts Reported ±1 SE ranges do not overlap
Code Migration 34.80% ±4.22 44.82% ±4.21 10.02 pts Reported ±1 SE ranges do not overlap
IOI 67.83% ±3.61 57.72% ±1.99 10.11 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 6.06% ±0.88 28.28% ±1.34 22.22 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 70.39% ±4.84 86.17% ±2.18 15.78 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 43.75% ±2.21 69.46% ±5.67 25.71 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Gemini 3.7 Flash average Grok 4.7 average
Index 38.79% 40.21%
Legal 43.54% 48.00%
Finance 62.68% 61.61%
Healthcare 68.67% 69.47%
Math 58.00% 26.00%
Science 5.71% 11.43%
Education 49.23% 40.79%
Coding 35.82% 43.50%
Beta 43.75% 69.46%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.7 Flash cost Grok 4.7 cost Gemini 3.7 Flash latency Grok 4.7 latency
Vals Index $4.17 $11.16 43m13s 32m56s
Vals RSI Index N/A N/A 90h00m 90h00m
Harvey's Legal Agent Benchmark $2.56 $11.13 7m32s 42m05s
Legal Research Bench $0.76 $4.92 4m49s 21m14s
LegalBench N/A N/A 2.23s 26.02s
EMB $5.69 $6.48 10m51s 34m57s
Finance Agent (v2) $1.48 $2.72 3m01s 15m33s
Tax Agent Bench $0.48 $1.80 87.24s 16m36s
MedCode N/A N/A 9.16s 3m33s
MedScribe N/A N/A 16.82s 2m33s
ProofBench v1.1 $0.56 $0.79 7m04s 12m47s
Terminal-Bench Science $13.22 $21.40 1h08m 2h40m
SAGE N/A N/A 9.79s 4m51s
Code Migration $21.46 $36.55 2h39m 1h09m
IOI $3.40 $12.71 21m09s 46m48s
ProgramBench $7.13 $46.49 21m30s 3h38m
Terminal-Bench 4.0 $14.07 $14.32 1h59m 1h40m
Vibe Code Bench v1.1 $4.83 $15.83 23m43s 35m47s
CyberBench v1.1 $2.15 $8.54 7m34s 15m30s

Results available only for Gemini 3.7 Flash

  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • SRE Bench

Results available only for Grok 4.7

  • BioMysteryBench
  • MysteryMechanism
  • Public Benefits Bench v1.1
Model details Gemini 3.7 Flash Model details Grok 4.7