Model comparison

Claude Sonnet 5 vs Gemini 3.7 Flash: Benchmark Comparison

Claude Sonnet 5 has the higher score on 10 of 26 shared benchmarks; Gemini 3.7 Flash leads on 15.

The largest observed score gap is 22.83 pts on IOI , where Gemini 3.7 Flash leads.

Reported ±1 standard-error ranges overlap on 10 of 26 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 Gemini 3.7 Flash Gap Reported uncertainty
Vals Index 59.61% ±1.15 59.31% ±1.06 0.30 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 8.75% ±2.15 3.75 pts Reported ±1 SE ranges overlap
Legal Research Bench 41.83% ±3.43 34.62% ±3.31 7.21 pts Reported ±1 SE ranges do not overlap
LegalBench 83.92% ±0.46 87.26% ±0.42 3.33 pts Reported ±1 SE ranges do not overlap
EMB 66.32% ±3.01 71.33% ±2.26 5.01 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 59.04% ±0.27 5.13 pts Reported ±1 SE ranges do not overlap
MortgageTax 70.03% ±0.90 66.65% ±0.92 3.38 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 57.66% ±3.31 4.61 pts Reported ±1 SE ranges overlap
TaxEval v2 75.63% ±0.84 74.73% ±0.85 0.90 pts Reported ±1 SE ranges overlap
MedCode 47.54% ±2.27 53.39% ±2.12 5.85 pts Reported ±1 SE ranges do not overlap
MedScribe 76.05% ±3.05 83.94% ±2.00 7.89 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 58.00% ±4.96 19.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 2.86% ±2.01 5.71% ±2.79 2.86 pts Reported ±1 SE ranges overlap
GPQA Diamond 88.89% ±2.22 93.94% ±1.49 5.05 pts Reported ±1 SE ranges do not overlap
MMLU Pro 87.55% ±0.37 90.12% ±0.30 2.58 pts Reported ±1 SE ranges do not overlap
MMMU Pro 83.01% ±0.90 88.96% ±0.75 5.95 pts Reported ±1 SE ranges do not overlap
SAGE 48.92% ±3.40 49.23% ±3.38 0.31 pts Reported ±1 SE ranges overlap
Code Migration 44.39% ±4.25 34.80% ±4.22 9.59 pts Reported ±1 SE ranges do not overlap
IOI 45.00% ±2.75 67.83% ±3.61 22.83 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 82.43% ±1.09 88.65% ±0.92 6.22 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SkillsBench 46.48% ±4.49 65.89% ±4.55 19.42 pts Reported ±1 SE ranges do not overlap
SWE-bench 79.60% ±1.80 80.80% ±1.76 1.20 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 8.08% ±1.82 6.06% ±0.88 2.02 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 81.33% ±3.05 70.39% ±4.84 10.93 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 61.91% ±5.74 43.75% ±2.21 18.16 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Sonnet 5 average Gemini 3.7 Flash average
Index 59.61% 59.31%
Legal 43.58% 43.54%
Finance 65.63% 65.88%
Healthcare 61.80% 68.67%
Math 77.00% 58.00%
Science 2.86% 5.71%
Academic 86.48% 91.01%
Education 48.92% 49.23%
Coding 48.41% 51.80%
Beta 61.91% 43.75%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost Gemini 3.7 Flash cost Claude Sonnet 5 latency Gemini 3.7 Flash latency
Vals Index $11.74 $4.17 45m46s 43m13s
Harvey's Legal Agent Benchmark $8.95 $2.56 38m52s 7m32s
Legal Research Bench $2.72 $0.76 25m46s 4m49s
LegalBench N/A N/A 4.90s 2.23s
EMB $10.29 $5.69 49m46s 10m51s
Finance Agent (v2) $0.75 $1.48 13m12s 3m01s
MortgageTax N/A N/A 28.67s 4.11s
Tax Agent Bench $1.79 $0.48 19m19s 87.24s
TaxEval v2 N/A N/A 3m22s 7.83s
MedCode N/A N/A 2m15s 9.16s
MedScribe N/A N/A 4m12s 16.82s
ProofBench v1.1 $1.37 $0.56 14m58s 7m04s
Terminal-Bench Science $18.70 $13.22 2h55m 1h08m
GPQA Diamond N/A N/A 63.23s 8.77s
MMLU Pro N/A N/A 25.19s 4.16s
MMMU Pro N/A N/A 18.82s 6.28s
SAGE N/A N/A 7m14s 9.79s
Code Migration $35.31 $21.46 1h57m 2h39m
IOI $12.89 $3.40 59m44s 21m09s
LiveCodeBench N/A N/A 77.02s 14.53s
ProgramBench $24.36 $7.13 1h31m 21m30s
SkillsBench $2.90 $1.80 15m12s 3m06s
SWE-bench $1.49 $1.44 16m02s 5m50s
Terminal-Bench 4.0 $14.25 $14.07 3h16m 1h59m
Vibe Code Bench v1.1 $25.39 $4.83 1h07m 23m43s
CyberBench v1.1 $1.89 $2.15 16m57s 7m34s

Results available only for Claude Sonnet 5

  • Vibe Code Bench 1-100
  • Public Benefits Bench v1.1

Results available only for Gemini 3.7 Flash

  • Vals RSI Index
  • SRE Bench
Model details Claude Sonnet 5 Model details Gemini 3.7 Flash