Gemini 3.8 Flash vs GPT-5.6 Sol: Benchmark Comparison

Gemini 3.8 Flash has the higher score on 9 of 33 shared benchmarks; GPT-5.6 Sol leads on 24.

The largest observed score gap is 35.00 pts on ProofBench v1.1 , where GPT-5.6 Sol leads.

Reported ±1 standard-error ranges overlap on 14 of 31 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.8 Flash GPT-5.6 Sol Gap Reported uncertainty
Vals Index 54.83% ±1.03 58.01% ±1.02 3.18 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 19.93% 23.88% 3.95 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 10.00% ±2.42 2.50% ±0.83 7.50 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 38.94% ±3.39 48.08% ±3.47 9.13 pts Reported ±1 SE ranges do not overlap
LegalBench 86.99% ±0.43 86.97% ±0.41 0.03 pts Reported ±1 SE ranges overlap
EMB 72.20% ±2.42 72.34% ±2.37 0.14 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 61.44% ±0.13 53.76% ±0.85 7.68 pts Reported ±1 SE ranges do not overlap
MortgageTax 65.34% ±0.94 67.29% ±0.92 1.95 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 66.77% ±3.15 67.95% ±3.12 1.18 pts Reported ±1 SE ranges overlap
TaxEval v2 74.45% ±0.85 74.78% ±0.86 0.33 pts Reported ±1 SE ranges overlap
MedCode 48.13% ±2.18 43.97% ±2.26 4.16 pts Reported ±1 SE ranges overlap
MedScribe 84.50% ±1.94 85.23% ±1.97 0.74 pts Reported ±1 SE ranges overlap
ProofBench v1.1 48.00% ±5.02 83.00% ±3.77 35.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 62.22% ±1.70 71.11% ±0.64 8.89 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 36.49% ±3.24 33.33% ±3.17 3.15 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 8.57% ±3.37 20.00% ±4.82 11.43 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 94.44% ±1.48 95.20% ±1.07 0.76 pts Reported ±1 SE ranges overlap
MMLU Pro 90.22% ±0.29 89.10% ±0.31 1.12 pts Reported ±1 SE ranges do not overlap
MMMU Pro 89.08% ±0.75 88.84% ±0.76 0.23 pts Reported ±1 SE ranges overlap
SAGE 35.06% ±3.36 52.56% ±3.42 17.50 pts Reported ±1 SE ranges do not overlap
Code Migration 36.55% ±4.18 52.92% ±4.35 16.37 pts Reported ±1 SE ranges do not overlap
IOI 56.94% ±2.26 91.17% ±4.51 34.22 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 89.48% ±0.90 82.60% ±1.09 6.88 pts Reported ±1 SE ranges do not overlap
ProgramBench 1.00% ±0.70 1.50% ±0.86 0.50 pts Reported ±1 SE ranges overlap
SkillsBench 57.98% ±4.42 54.10% ±4.73 3.88 pts Reported ±1 SE ranges overlap
SWE-bench 80.00% ±1.79 96.20% ±0.86 16.20 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 19.19% ±2.52 37.88% ±0.00 18.69 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 18.77% ±3.31 20.01% ±3.63 1.24 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 78.65% ±3.88 80.50% ±3.72 1.84 pts Reported ±1 SE ranges overlap
CyberBench v1.1 43.75% ±2.21 76.31% ±5.18 32.56 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 65.29% ±1.24 66.51% ±1.23 1.22 pts Reported ±1 SE ranges overlap
CUA-bench 4.17% 8.33% 4.17 pts Uncertainty comparison unavailable
Time Horizon Index: KSP 18.83% ±0.00 23.83% ±0.00 5.00 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Gemini 3.8 Flash average GPT-5.6 Sol average
Legal 45.31% 45.85%
Finance 68.04% 67.22%
Healthcare 66.32% 64.60%
Math 48.00% 83.00%
Science 35.76% 41.48%
Academic 91.25% 91.05%
Education 35.06% 52.56%
Coding 48.73% 57.43%
Cyber 43.75% 76.31%
Social Mobility 65.29% 66.51%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.8 Flash cost GPT-5.6 Sol cost Gemini 3.8 Flash latency GPT-5.6 Sol latency
Vals Index $5.73 $14.24 51m57s 37m23s
Vals RSI Index $332.04 $347.10 90h00m 90h00m
Harvey's Legal Agent Benchmark $3.66 $10.37 29m21s 22m50s
Legal Research Bench $1.63 $21.61 5m24s 1h17m
LegalBench N/A N/A 3.32s 6.20s
EMB $8.24 $6.01 12m49s 9m50s
Finance Agent (v2) $2.00 $1.25 3m21s 19m25s
MortgageTax N/A N/A 14.79s 12.62s
Tax Agent Bench $0.86 $7.51 2m58s 44m41s
TaxEval v2 N/A N/A 8.76s 62.03s
MedCode N/A N/A 42.89s 96.59s
MedScribe N/A N/A 21.14s 94.40s
ProofBench v1.1 $0.60 $1.44 6m09s 8m40s
BioMysteryBench $1.55 $1.93 5m27s 18m18s
MysteryMechanism $1.73 $0.96 5m22s 8m43s
Terminal-Bench Science $5.64 $7.32 55m11s 1h44m
GPQA Diamond N/A N/A 20.18s 58.30s
MMLU Pro N/A N/A 7.82s 15.65s
MMMU Pro N/A N/A 12.03s 26.20s
SAGE N/A N/A 28.64s 75.39s
Code Migration $18.49 $24.54 2h37m 59m12s
IOI $3.98 $7.98 14m20s 1h02m
LiveCodeBench N/A N/A 24.07s 56.51s
ProgramBench $10.93 $15.29 46m39s 29m24s
SkillsBench $2.34 $4.14 3m55s 13m07s
SWE-bench $2.19 $1.15 11m23s 3m02s
Terminal-Bench 4.0 $8.77 $7.98 1h48m 33m46s
Vibe Code Bench 1-100 $18.91 $47.43 1h06m 2h37m
Vibe Code Bench v1.1 $6.87 $33.40 8m39s 33m53s
CyberBench v1.1 $1.31 $3.05 9m49s 10m47s
Public Benefits Bench v1.1 $1.01 $8.85 3m21s 1h01m
CUA-bench $280.03 $206.99 36.55s 18.80s
Time Horizon Index: KSP $543.14 $2027.20 N/A N/A

Results available only for Gemini 3.8 Flash

None.

Results available only for GPT-5.6 Sol

  • SRE Bench
Model details Gemini 3.8 Flash Model details GPT-5.6 Sol