Claude Opus 5 vs Gemini 3.8 Flash: Benchmark Comparison

Claude Opus 5 has the higher score on 27 of 33 shared benchmarks; Gemini 3.8 Flash leads on 5.

The largest observed score gap is 51.00 pts on ProofBench v1.1 , where Claude Opus 5 leads.

Reported ±1 standard-error ranges overlap on 10 of 31 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5 Gemini 3.8 Flash Gap Reported uncertainty
Vals Index 63.67% ±0.96 54.83% ±1.03 8.85 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 33.02% 19.93% 13.09 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 6.67% ±1.65 10.00% ±2.42 3.33 pts Reported ±1 SE ranges overlap
Legal Research Bench 55.29% ±3.46 38.94% ±3.39 16.35 pts Reported ±1 SE ranges do not overlap
LegalBench 86.97% ±0.42 86.99% ±0.43 0.02 pts Reported ±1 SE ranges overlap
EMB 73.56% ±2.24 72.20% ±2.42 1.36 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 58.63% ±0.08 61.44% ±0.13 2.80 pts Reported ±1 SE ranges do not overlap
MortgageTax 72.06% ±0.88 65.34% ±0.94 6.72 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 75.06% ±2.96 66.77% ±3.15 8.29 pts Reported ±1 SE ranges do not overlap
TaxEval v2 75.14% ±0.83 74.45% ±0.85 0.70 pts Reported ±1 SE ranges overlap
MedCode 63.57% ±1.99 48.13% ±2.18 15.44 pts Reported ±1 SE ranges do not overlap
MedScribe 90.98% ±1.92 84.50% ±1.94 6.49 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 99.00% ±1.00 48.00% ±5.02 51.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 79.26% ±2.06 62.22% ±1.70 17.04 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 37.39% ±3.25 36.49% ±3.24 0.90 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 27.14% ±5.35 8.57% ±3.37 18.57 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 93.43% ±1.24 94.44% ±1.48 1.01 pts Reported ±1 SE ranges overlap
MMLU Pro 91.59% ±0.28 90.22% ±0.29 1.37 pts Reported ±1 SE ranges do not overlap
MMMU Pro 89.88% ±0.72 89.08% ±0.75 0.81 pts Reported ±1 SE ranges overlap
SAGE 49.43% ±3.29 35.06% ±3.36 14.36 pts Reported ±1 SE ranges do not overlap
Code Migration 57.47% ±4.37 36.55% ±4.18 20.93 pts Reported ±1 SE ranges do not overlap
IOI 84.33% ±9.96 56.94% ±2.26 27.39 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 89.03% ±0.91 89.48% ±0.90 0.45 pts Reported ±1 SE ranges overlap
ProgramBench 3.00% ±1.21 1.00% ±0.70 2.00 pts Reported ±1 SE ranges do not overlap
SkillsBench 60.44% ±4.58 57.98% ±4.42 2.46 pts Reported ±1 SE ranges overlap
SWE-bench 97.00% ±0.76 80.00% ±1.79 17.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 53.53% ±1.34 19.19% ±2.52 34.34 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 28.53% ±4.33 18.77% ±3.31 9.76 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 88.40% ±3.00 78.65% ±3.88 9.75 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 65.36% ±5.55 43.75% ±2.21 21.61 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 76.93% ±1.10 65.29% ±1.24 11.64 pts Reported ±1 SE ranges do not overlap
CUA-bench 9.00% 4.17% 4.83 pts Uncertainty comparison unavailable
Time Horizon Index: KSP 18.83% ±0.00 18.83% ±0.00 0.00 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Opus 5 average Gemini 3.8 Flash average
Legal 49.64% 45.31%
Finance 70.89% 68.04%
Healthcare 77.28% 66.32%
Math 99.00% 48.00%
Science 47.93% 35.76%
Academic 91.64% 91.25%
Education 49.43% 35.06%
Coding 62.42% 48.73%
Cyber 65.36% 43.75%
Social Mobility 76.93% 65.29%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5 cost Gemini 3.8 Flash cost Claude Opus 5 latency Gemini 3.8 Flash latency
Vals Index $19.31 $5.73 57m41s 51m57s
Vals RSI Index $418.69 $332.04 90h00m 90h00m
Harvey's Legal Agent Benchmark $23.67 $3.66 55m37s 29m21s
Legal Research Bench $6.76 $1.63 31m01s 5m24s
LegalBench N/A N/A 4.42s 3.32s
EMB $6.07 $8.24 25m59s 12m49s
Finance Agent (v2) $5.12 $2.00 9m58s 3m21s
MortgageTax N/A N/A 7.98s 14.79s
Tax Agent Bench $5.14 $0.86 16m48s 2m58s
TaxEval v2 N/A N/A 58.34s 8.76s
MedCode N/A N/A 24.50s 42.89s
MedScribe N/A N/A 76.56s 21.14s
ProofBench v1.1 $2.11 $0.60 9m51s 6m09s
BioMysteryBench $3.29 $1.55 36m36s 5m27s
MysteryMechanism $3.28 $1.73 16m53s 5m22s
Terminal-Bench Science $32.54 $5.64 2h12m 55m11s
GPQA Diamond N/A N/A 39.59s 20.18s
MMLU Pro N/A N/A 9.60s 7.82s
MMMU Pro N/A N/A 30.03s 12.03s
SAGE N/A N/A 56.17s 28.64s
Code Migration $60.51 $18.49 2h51m 2h37m
IOI $16.48 $3.98 56m42s 14m20s
LiveCodeBench N/A N/A 59.22s 24.07s
ProgramBench $60.29 $10.93 2h55m 46m39s
SkillsBench $2.51 $2.34 8m25s 3m55s
SWE-bench $1.29 $2.19 9m37s 11m23s
Terminal-Bench 4.0 $18.60 $8.77 1h07m 1h48m
Vibe Code Bench 1-100 $41.49 $18.91 1h22m 1h06m
Vibe Code Bench v1.1 $33.88 $6.87 1h28m 8m39s
CyberBench v1.1 $2.87 $1.31 19m25s 9m49s
Public Benefits Bench v1.1 $3.95 $1.01 30m09s 3m21s
CUA-bench $232.23 $280.03 25.29s 36.55s
Time Horizon Index: KSP $1348.77 $543.14 N/A N/A

Results available only for Claude Opus 5

  • SRE Bench

Results available only for Gemini 3.8 Flash

None.

Model details Claude Opus 5 Model details Gemini 3.8 Flash