Model comparison

Gemini 3.7 Flash vs GPT 5.5: Benchmark Comparison

Gemini 3.7 Flash has the higher score on 14 of 23 shared benchmarks; GPT 5.5 leads on 9.

The largest observed score gap is 10.36 pts on Code Migration , where GPT 5.5 leads.

Reported ±1 standard-error ranges overlap on 15 of 22 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.7 Flash GPT 5.5 Gap Reported uncertainty
Vals Index 59.31% ±1.06 57.41% ±1.15 1.90 pts Reported ±1 SE ranges overlap
Vals RSI Index 18.27% 16.48% 1.79 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 8.75% ±2.15 3.75% ±1.17 5.00 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 34.62% ±3.31 40.38% ±3.41 5.77 pts Reported ±1 SE ranges overlap
LegalBench 87.26% ±0.42 86.52% ±0.41 0.74 pts Reported ±1 SE ranges overlap
EMB 71.33% ±2.26 64.54% ±2.87 6.78 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 59.04% ±0.27 51.76% ±0.55 7.28 pts Reported ±1 SE ranges do not overlap
MortgageTax 66.65% ±0.92 68.76% ±0.91 2.11 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 57.66% ±3.31 60.46% ±3.22 2.79 pts Reported ±1 SE ranges overlap
TaxEval v2 74.73% ±0.85 74.98% ±0.85 0.25 pts Reported ±1 SE ranges overlap
MedCode 53.39% ±2.12 49.10% ±2.19 4.29 pts Reported ±1 SE ranges overlap
MedScribe 83.94% ±2.00 86.87% ±1.93 2.93 pts Reported ±1 SE ranges overlap
GPQA Diamond 93.94% ±1.49 93.18% ±1.29 0.76 pts Reported ±1 SE ranges overlap
MMLU Pro 90.12% ±0.30 88.14% ±0.32 1.98 pts Reported ±1 SE ranges do not overlap
MMMU Pro 88.96% ±0.75 88.27% ±0.77 0.69 pts Reported ±1 SE ranges overlap
SAGE 49.23% ±3.38 51.53% ±3.95 2.30 pts Reported ±1 SE ranges overlap
Code Migration 34.80% ±4.22 45.16% ±4.16 10.36 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 88.65% ±0.92 85.30% ±1.02 3.36 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
SkillsBench 65.89% ±4.55 62.21% ±4.52 3.68 pts Reported ±1 SE ranges overlap
SWE-bench 80.80% ±1.76 82.60% ±1.70 1.80 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 70.39% ±4.84 69.85% ±4.54 0.55 pts Reported ±1 SE ranges overlap
SRE Bench 4.58% ±1.29 3.82% ±1.19 0.76 pts Reported ±1 SE ranges overlap

Performance by category

Category Gemini 3.7 Flash average GPT 5.5 average
Index 38.79% 36.95%
Legal 43.54% 43.55%
Finance 65.88% 64.10%
Healthcare 68.67% 67.98%
Academic 91.01% 89.86%
Education 49.23% 51.53%
Coding 56.76% 57.60%
Beta 4.58% 3.82%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.7 Flash cost GPT 5.5 cost Gemini 3.7 Flash latency GPT 5.5 latency
Vals Index $4.17 $6.14 43m13s 20m51s
Vals RSI Index $180.16 $857.16 90h00m 90h00m
Harvey's Legal Agent Benchmark $2.56 $4.60 7m32s 12m14s
Legal Research Bench $0.76 $7.40 4m49s 37m45s
LegalBench N/A N/A 2.23s 18.14s
EMB $5.69 $3.27 10m51s 15m11s
Finance Agent (v2) $1.48 $4.15 3m01s 11m02s
MortgageTax N/A N/A 4.11s 28.24s
Tax Agent Bench $0.48 $4.07 87.24s 15m52s
TaxEval v2 N/A N/A 7.83s 95.83s
MedCode N/A N/A 9.16s 2m40s
MedScribe N/A N/A 16.82s 2m13s
GPQA Diamond N/A N/A 8.77s 112.33s
MMLU Pro N/A N/A 4.16s 42.11s
MMMU Pro N/A N/A 6.28s 54.15s
SAGE N/A N/A 9.79s 76.05s
Code Migration $21.46 $6.44 2h39m 31m35s
LiveCodeBench N/A N/A 14.53s 2m47s
ProgramBench $7.13 $6.95 21m30s 22m03s
SkillsBench $1.80 $2.54 3m06s 6m55s
SWE-bench $1.44 $1.36 5m50s 7m06s
Vibe Code Bench v1.1 $4.83 $16.66 23m43s 31m52s
SRE Bench $21.33 $17.77 1h50m 46m28s

Results available only for Gemini 3.7 Flash

  • ProofBench v1.1
  • Terminal-Bench Science
  • IOI
  • Terminal-Bench 4.0
  • CyberBench v1.1

Results available only for GPT 5.5

  • Time Horizon Index: KSP
  • Public Benefits Bench v1.1
Model details Gemini 3.7 Flash Model details GPT 5.5