Gemini 3.1 Pro Preview (02/26) vs GPT 5.4 Mini: Benchmark Comparison

Gemini 3.1 Pro Preview (02/26) has the higher score on 12 of 18 shared benchmarks; GPT 5.4 Mini leads on 3.

The largest observed score gap is 15.94 pts on Vibe Code Bench v1.1 , where GPT 5.4 Mini leads.

Reported ±1 standard-error ranges overlap on 8 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Gemini 3.1 Pro Preview (02/26) GPT 5.4 Mini Gap Reported uncertainty
Vals Index 33.44% ±1.13 33.17% ±1.28 0.27 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
Legal Research Bench 20.67% ±2.81 12.50% ±2.30 8.17 pts Reported ±1 SE ranges do not overlap
EMB 52.62% ±2.97 45.43% ±3.65 7.19 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 42.98% ±1.21 45.36% ±0.45 2.38 pts Reported ±1 SE ranges do not overlap
MortgageTax 69.40% ±0.91 63.51% ±0.91 5.88 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 37.84% ±3.10 36.59% ±3.10 1.26 pts Reported ±1 SE ranges overlap
TaxEval v2 72.88% ±0.86 71.22% ±0.90 1.66 pts Reported ±1 SE ranges overlap
GPQA Diamond 95.45% ±1.05 83.08% ±2.46 12.37 pts Reported ±1 SE ranges do not overlap
MMLU Pro 90.99% ±0.28 84.55% ±0.36 6.43 pts Reported ±1 SE ranges do not overlap
MMMU Pro 88.21% ±0.78 79.25% ±0.97 8.96 pts Reported ±1 SE ranges do not overlap
SAGE 48.68% ±3.29 50.81% ±3.40 2.14 pts Reported ±1 SE ranges overlap
Code Migration 17.31% ±3.93 12.94% ±3.64 4.38 pts Reported ±1 SE ranges overlap
LiveCodeBench 88.48% ±0.93 81.47% ±1.09 7.02 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SWE-bench 78.80% ±1.83 73.00% ±1.99 5.80 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 2.52% ±0.51 2.52% ±1.01 0.00 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 32.03% ±4.34 47.97% ±5.61 15.94 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Gemini 3.1 Pro Preview (02/26) average GPT 5.4 Mini average
Legal 10.34% 6.25%
Finance 55.14% 52.42%
Academic 91.55% 82.29%
Education 48.68% 50.81%
Coding 36.53% 36.32%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Gemini 3.1 Pro Preview (02/26) cost GPT 5.4 Mini cost Gemini 3.1 Pro Preview (02/26) latency GPT 5.4 Mini latency
Vals Index $2.18 $1.45 9m52s 29m49s
Harvey's Legal Agent Benchmark $1.15 $1.00 7m17s 11m21s
Legal Research Bench $1.10 $1.86 4m02s 37m36s
EMB $3.77 $2.15 11m49s 43m21s
Finance Agent (v2) $1.59 $1.20 3m28s 35m54s
MortgageTax N/A N/A 23.62s 2m06s
Tax Agent Bench $0.30 $0.81 100.96s 16m46s
TaxEval v2 N/A N/A 41.93s 44.58s
GPQA Diamond N/A N/A 65.76s 63.11s
MMLU Pro N/A N/A 23.91s 20.00s
MMMU Pro N/A N/A 76.99s 50.81s
SAGE N/A N/A 61.17s 2m21s
Code Migration $1.61 $1.95 12m01s 28m53s
LiveCodeBench N/A N/A 88.89s 3m09s
ProgramBench N/A N/A 12m36s 53m45s
SWE-bench $0.78 $0.51 5m12s 5m26s
Terminal-Bench 4.0 $4.11 $1.43 18m44s 29m41s
Vibe Code Bench v1.1 $3.83 $1.19 20m12s 34m15s

Results available only for Gemini 3.1 Pro Preview (02/26)

  • LegalBench
  • MedCode
  • MedScribe
  • ProofBench v1.1
  • BioMysteryBench
  • MysteryMechanism
  • Terminal-Bench Science
  • IOI
  • Vibe Code Bench 1-100
  • Public Benefits Bench v1.1

Results available only for GPT 5.4 Mini

None.

Model details Gemini 3.1 Pro Preview (02/26) Model details GPT 5.4 Mini