Gemini 3.6 Flash vs GPT 5.5: Benchmark Comparison

Gemini 3.6 Flash has the higher score on 8 of 20 shared benchmarks; GPT 5.5 leads on 12.

The largest observed score gap is 15.38 pts on Legal Research Bench , where GPT 5.5 leads.

Reported ±1 standard-error ranges overlap on 12 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.6 Flash GPT 5.5 Gap Reported uncertainty
Harvey's Legal Agent Benchmark 3.33% ±1.17 3.75% ±1.17 0.42 pts Reported ±1 SE ranges overlap
Legal Research Bench 25.00% ±3.01 40.38% ±3.41 15.38 pts Reported ±1 SE ranges do not overlap
LegalBench 86.70% ±0.41 86.52% ±0.41 0.19 pts Reported ±1 SE ranges overlap
EMB 65.41% ±2.65 64.54% ±2.87 0.87 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 56.30% ±0.18 51.76% ±0.55 4.54 pts Reported ±1 SE ranges do not overlap
MortgageTax 67.61% ±0.92 68.76% ±0.91 1.15 pts Reported ±1 SE ranges overlap
Tax Agent Bench 49.90% ±3.30 60.46% ±3.22 10.56 pts Reported ±1 SE ranges do not overlap
TaxEval v2 74.86% ±0.85 74.98% ±0.85 0.12 pts Reported ±1 SE ranges overlap
MedCode 53.15% ±2.16 49.10% ±2.19 4.05 pts Reported ±1 SE ranges overlap
MedScribe 79.66% ±1.86 86.87% ±1.93 7.21 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 93.43% ±1.33 93.18% ±1.29 0.25 pts Reported ±1 SE ranges overlap
MMLU Pro 89.28% ±0.30 88.14% ±0.32 1.13 pts Reported ±1 SE ranges do not overlap
MMMU Pro 88.38% ±0.77 88.27% ±0.77 0.12 pts Reported ±1 SE ranges overlap
SAGE 48.69% ±3.38 51.53% ±3.95 2.85 pts Reported ±1 SE ranges overlap
Code Migration 30.93% ±4.07 45.16% ±4.16 14.23 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 88.08% ±0.94 85.30% ±1.02 2.78 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
SWE-bench 79.60% ±1.80 82.60% ±1.70 3.00 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 64.00% ±4.29 69.85% ±4.54 5.84 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 56.56% ±1.29 60.89% ±1.27 4.33 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Gemini 3.6 Flash average GPT 5.5 average
Legal 38.35% 43.55%
Finance 62.81% 64.10%
Healthcare 66.41% 67.98%
Academic 90.36% 89.86%
Education 48.69% 51.53%
Coding 52.52% 56.68%
Social Mobility 56.56% 60.89%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.6 Flash cost GPT 5.5 cost Gemini 3.6 Flash latency GPT 5.5 latency
Harvey's Legal Agent Benchmark $1.90 $4.60 16m24s 12m14s
Legal Research Bench $0.47 $7.40 2m48s 37m45s
LegalBench N/A N/A 2.96s 18.14s
EMB $3.89 $3.27 12m06s 15m11s
Finance Agent (v2) $1.40 $4.15 3m45s 11m02s
MortgageTax N/A N/A 9.24s 28.24s
Tax Agent Bench $0.36 $4.07 105.32s 15m52s
TaxEval v2 N/A N/A 12.34s 95.83s
MedCode N/A N/A 16.68s 2m40s
MedScribe N/A N/A 33.55s 2m13s
GPQA Diamond N/A N/A 18.06s 112.33s
MMLU Pro N/A N/A 7.30s 42.11s
MMMU Pro N/A N/A 11.31s 54.15s
SAGE N/A N/A 21.96s 76.05s
Code Migration $8.93 $6.44 1h30m 31m35s
LiveCodeBench N/A N/A 31.88s 2m47s
ProgramBench $5.92 $6.95 47m33s 22m03s
SWE-bench $1.19 $1.36 13m08s 7m06s
Vibe Code Bench v1.1 $3.04 $16.66 25m23s 31m52s
Public Benefits Bench v1.1 $0.39 $3.97 2m02s 39m37s

Results available only for Gemini 3.6 Flash

  • BioMysteryBench
  • Terminal-Bench Science
  • IOI
  • CyberBench v1.1

Results available only for GPT 5.5

  • Vals RSI Index
  • SkillsBench
  • SRE Bench
  • Time Horizon Index: KSP
Model details Gemini 3.6 Flash Model details GPT 5.5