Claude Opus 5 vs Gemini 3.8 Flash: Benchmark Comparison
Claude Opus 5 has the higher score on 27 of 33 shared benchmarks; Gemini 3.8 Flash leads on 5.
The largest observed score gap is 51.00 pts on ProofBench v1.1 , where Claude Opus 5 leads.
Reported ±1 standard-error ranges overlap on 10 of 31 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Claude Opus 5 | Gemini 3.8 Flash | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 63.67% ±0.96 | 54.83% ±1.03 | 8.85 pts | Reported ±1 SE ranges do not overlap |
| Vals RSI Index | 33.02% | 19.93% | 13.09 pts | Uncertainty comparison unavailable |
| Harvey's Legal Agent Benchmark | 6.67% ±1.65 | 10.00% ±2.42 | 3.33 pts | Reported ±1 SE ranges overlap |
| Legal Research Bench | 55.29% ±3.46 | 38.94% ±3.39 | 16.35 pts | Reported ±1 SE ranges do not overlap |
| LegalBench | 86.97% ±0.42 | 86.99% ±0.43 | 0.02 pts | Reported ±1 SE ranges overlap |
| EMB | 73.56% ±2.24 | 72.20% ±2.42 | 1.36 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 58.63% ±0.08 | 61.44% ±0.13 | 2.80 pts | Reported ±1 SE ranges do not overlap |
| MortgageTax | 72.06% ±0.88 | 65.34% ±0.94 | 6.72 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 75.06% ±2.96 | 66.77% ±3.15 | 8.29 pts | Reported ±1 SE ranges do not overlap |
| TaxEval v2 | 75.14% ±0.83 | 74.45% ±0.85 | 0.70 pts | Reported ±1 SE ranges overlap |
| MedCode | 63.57% ±1.99 | 48.13% ±2.18 | 15.44 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 90.98% ±1.92 | 84.50% ±1.94 | 6.49 pts | Reported ±1 SE ranges do not overlap |
| ProofBench v1.1 | 99.00% ±1.00 | 48.00% ±5.02 | 51.00 pts | Reported ±1 SE ranges do not overlap |
| BioMysteryBench | 79.26% ±2.06 | 62.22% ±1.70 | 17.04 pts | Reported ±1 SE ranges do not overlap |
| MysteryMechanism | 37.39% ±3.25 | 36.49% ±3.24 | 0.90 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench Science | 27.14% ±5.35 | 8.57% ±3.37 | 18.57 pts | Reported ±1 SE ranges do not overlap |
| GPQA Diamond | 93.43% ±1.24 | 94.44% ±1.48 | 1.01 pts | Reported ±1 SE ranges overlap |
| MMLU Pro | 91.59% ±0.28 | 90.22% ±0.29 | 1.37 pts | Reported ±1 SE ranges do not overlap |
| MMMU Pro | 89.88% ±0.72 | 89.08% ±0.75 | 0.81 pts | Reported ±1 SE ranges overlap |
| SAGE | 49.43% ±3.29 | 35.06% ±3.36 | 14.36 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 57.47% ±4.37 | 36.55% ±4.18 | 20.93 pts | Reported ±1 SE ranges do not overlap |
| IOI | 84.33% ±9.96 | 56.94% ±2.26 | 27.39 pts | Reported ±1 SE ranges do not overlap |
| LiveCodeBench | 89.03% ±0.91 | 89.48% ±0.90 | 0.45 pts | Reported ±1 SE ranges overlap |
| ProgramBench | 3.00% ±1.21 | 1.00% ±0.70 | 2.00 pts | Reported ±1 SE ranges do not overlap |
| SkillsBench | 60.44% ±4.58 | 57.98% ±4.42 | 2.46 pts | Reported ±1 SE ranges overlap |
| SWE-bench | 97.00% ±0.76 | 80.00% ±1.79 | 17.00 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench 4.0 | 53.53% ±1.34 | 19.19% ±2.52 | 34.34 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench 1-100 | 28.53% ±4.33 | 18.77% ±3.31 | 9.76 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 88.40% ±3.00 | 78.65% ±3.88 | 9.75 pts | Reported ±1 SE ranges do not overlap |
| CyberBench v1.1 | 65.36% ±5.55 | 43.75% ±2.21 | 21.61 pts | Reported ±1 SE ranges do not overlap |
| Public Benefits Bench v1.1 | 76.93% ±1.10 | 65.29% ±1.24 | 11.64 pts | Reported ±1 SE ranges do not overlap |
| CUA-bench | 9.00% | 4.17% | 4.83 pts | Uncertainty comparison unavailable |
| Time Horizon Index: KSP | 18.83% ±0.00 | 18.83% ±0.00 | 0.00 pts | Reported ±1 SE ranges overlap |
Performance by category
| Category | Claude Opus 5 average | Gemini 3.8 Flash average |
|---|---|---|
| Legal | 49.64% | 45.31% |
| Finance | 70.89% | 68.04% |
| Healthcare | 77.28% | 66.32% |
| Math | 99.00% | 48.00% |
| Science | 47.93% | 35.76% |
| Academic | 91.64% | 91.25% |
| Education | 49.43% | 35.06% |
| Coding | 62.42% | 48.73% |
| Cyber | 65.36% | 43.75% |
| Social Mobility | 76.93% | 65.29% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Claude Opus 5 cost | Gemini 3.8 Flash cost | Claude Opus 5 latency | Gemini 3.8 Flash latency |
|---|---|---|---|---|
| Vals Index | $19.31 | $5.73 | 57m41s | 51m57s |
| Vals RSI Index | $418.69 | $332.04 | 90h00m | 90h00m |
| Harvey's Legal Agent Benchmark | $23.67 | $3.66 | 55m37s | 29m21s |
| Legal Research Bench | $6.76 | $1.63 | 31m01s | 5m24s |
| LegalBench | N/A | N/A | 4.42s | 3.32s |
| EMB | $6.07 | $8.24 | 25m59s | 12m49s |
| Finance Agent (v2) | $5.12 | $2.00 | 9m58s | 3m21s |
| MortgageTax | N/A | N/A | 7.98s | 14.79s |
| Tax Agent Bench | $5.14 | $0.86 | 16m48s | 2m58s |
| TaxEval v2 | N/A | N/A | 58.34s | 8.76s |
| MedCode | N/A | N/A | 24.50s | 42.89s |
| MedScribe | N/A | N/A | 76.56s | 21.14s |
| ProofBench v1.1 | $2.11 | $0.60 | 9m51s | 6m09s |
| BioMysteryBench | $3.29 | $1.55 | 36m36s | 5m27s |
| MysteryMechanism | $3.28 | $1.73 | 16m53s | 5m22s |
| Terminal-Bench Science | $32.54 | $5.64 | 2h12m | 55m11s |
| GPQA Diamond | N/A | N/A | 39.59s | 20.18s |
| MMLU Pro | N/A | N/A | 9.60s | 7.82s |
| MMMU Pro | N/A | N/A | 30.03s | 12.03s |
| SAGE | N/A | N/A | 56.17s | 28.64s |
| Code Migration | $60.51 | $18.49 | 2h51m | 2h37m |
| IOI | $16.48 | $3.98 | 56m42s | 14m20s |
| LiveCodeBench | N/A | N/A | 59.22s | 24.07s |
| ProgramBench | $60.29 | $10.93 | 2h55m | 46m39s |
| SkillsBench | $2.51 | $2.34 | 8m25s | 3m55s |
| SWE-bench | $1.29 | $2.19 | 9m37s | 11m23s |
| Terminal-Bench 4.0 | $18.60 | $8.77 | 1h07m | 1h48m |
| Vibe Code Bench 1-100 | $41.49 | $18.91 | 1h22m | 1h06m |
| Vibe Code Bench v1.1 | $33.88 | $6.87 | 1h28m | 8m39s |
| CyberBench v1.1 | $2.87 | $1.31 | 19m25s | 9m49s |
| Public Benefits Bench v1.1 | $3.95 | $1.01 | 30m09s | 3m21s |
| CUA-bench | $232.23 | $280.03 | 25.29s | 36.55s |
| Time Horizon Index: KSP | $1348.77 | $543.14 | N/A | N/A |
Results available only for Claude Opus 5
- SRE Bench
Results available only for Gemini 3.8 Flash
None.