Claude Opus 5.5 vs Gemini 3.8 Flash: Benchmark Comparison
Claude Opus 5.5 has the higher score on 22 of 24 shared benchmarks; Gemini 3.8 Flash leads on 2.
The largest observed score gap is 72.50 pts on Time Horizon Index: KSP , where Claude Opus 5.5 leads.
Reported ±1 standard-error ranges overlap on 3 of 22 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Claude Opus 5.5 | Gemini 3.8 Flash | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 66.97% ±0.89 | 54.83% ±1.03 | 12.15 pts | Reported ±1 SE ranges do not overlap |
| Vals RSI Index | 37.31% | 19.93% | 17.38 pts | Uncertainty comparison unavailable |
| Harvey's Legal Agent Benchmark | 3.75% ±1.47 | 10.00% ±2.42 | 6.25 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 50.48% ±3.48 | 38.94% ±3.39 | 11.54 pts | Reported ±1 SE ranges do not overlap |
| EMB | 75.94% ±2.38 | 72.20% ±2.42 | 3.74 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 58.59% ±0.17 | 61.44% ±0.13 | 2.85 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 70.50% ±3.15 | 66.77% ±3.15 | 3.73 pts | Reported ±1 SE ranges overlap |
| MedCode | 49.80% ±2.27 | 48.13% ±2.18 | 1.66 pts | Reported ±1 SE ranges overlap |
| MedScribe | 91.43% ±1.93 | 84.50% ±1.94 | 6.93 pts | Reported ±1 SE ranges do not overlap |
| ProofBench v1.1 | 100.00% ±0.00 | 48.00% ±5.02 | 52.00 pts | Reported ±1 SE ranges do not overlap |
| BioMysteryBench | 79.26% ±0.98 | 62.22% ±1.70 | 17.04 pts | Reported ±1 SE ranges do not overlap |
| MysteryMechanism | 49.55% ±3.36 | 36.49% ±3.24 | 13.06 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 47.14% ±6.01 | 8.57% ±3.37 | 38.57 pts | Reported ±1 SE ranges do not overlap |
| SAGE | 45.83% ±3.36 | 35.06% ±3.36 | 10.77 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 66.65% ±4.33 | 36.55% ±4.18 | 30.10 pts | Reported ±1 SE ranges do not overlap |
| IOI | 95.06% ±4.94 | 56.94% ±2.26 | 38.11 pts | Reported ±1 SE ranges do not overlap |
| ProgramBench | 18.50% ±2.75 | 1.00% ±0.70 | 17.50 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench 4.0 | 65.15% ±0.00 | 19.19% ±2.52 | 45.96 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench 1-100 | 30.36% ±4.83 | 18.77% ±3.31 | 11.59 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 90.29% ±1.53 | 78.65% ±3.88 | 11.64 pts | Reported ±1 SE ranges do not overlap |
| CyberBench v1.1 | 55.36% ±5.13 | 43.75% ±2.21 | 11.61 pts | Reported ±1 SE ranges do not overlap |
| Public Benefits Bench v1.1 | 70.64% ±1.19 | 65.29% ±1.24 | 5.34 pts | Reported ±1 SE ranges do not overlap |
| CUA-bench | 14.00% | 4.17% | 9.83 pts | Uncertainty comparison unavailable |
| Time Horizon Index: KSP | 91.33% ±0.00 | 18.83% ±0.00 | 72.50 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Claude Opus 5.5 average | Gemini 3.8 Flash average |
|---|---|---|
| Legal | 27.12% | 24.47% |
| Finance | 68.34% | 66.80% |
| Healthcare | 70.61% | 66.32% |
| Math | 100.00% | 48.00% |
| Science | 58.65% | 35.76% |
| Education | 45.83% | 35.06% |
| Coding | 61.00% | 35.18% |
| Cyber | 55.36% | 43.75% |
| Social Mobility | 70.64% | 65.29% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Claude Opus 5.5 cost | Gemini 3.8 Flash cost | Claude Opus 5.5 latency | Gemini 3.8 Flash latency |
|---|---|---|---|---|
| Vals Index | $32.14 | $5.73 | 1h19m | 51m57s |
| Vals RSI Index | $411.86 | $332.04 | 90h00m | 90h00m |
| Harvey's Legal Agent Benchmark | $21.38 | $3.66 | 49m00s | 29m21s |
| Legal Research Bench | $25.16 | $1.63 | 1h52m | 5m24s |
| EMB | $11.02 | $8.24 | 35m06s | 12m49s |
| Finance Agent (v2) | $9.22 | $2.00 | 33m08s | 3m21s |
| Tax Agent Bench | $15.09 | $0.86 | 1h07m | 2m58s |
| MedCode | N/A | N/A | 4m07s | 42.89s |
| MedScribe | N/A | N/A | 7m22s | 21.14s |
| ProofBench v1.1 | $0.96 | $0.60 | 5m51s | 6m09s |
| BioMysteryBench | $3.64 | $1.55 | 12m19s | 5m27s |
| MysteryMechanism | $4.62 | $1.73 | 19m38s | 5m22s |
| Terminal-Bench Science | $19.12 | $5.64 | 2h28m | 55m11s |
| SAGE | N/A | N/A | 114.20s | 28.64s |
| Code Migration | $112.97 | $18.49 | 3h19m | 2h37m |
| IOI | $5.26 | $3.98 | 17m15s | 14m20s |
| ProgramBench | $68.05 | $10.93 | 2h24m | 46m39s |
| Terminal-Bench 4.0 | $13.20 | $8.77 | 1h04m | 1h48m |
| Vibe Code Bench 1-100 | $44.12 | $18.91 | 2h28m | 1h06m |
| Vibe Code Bench v1.1 | $57.92 | $6.87 | 1h36m | 8m39s |
| CyberBench v1.1 | $3.48 | $1.31 | 18m11s | 9m49s |
| Public Benefits Bench v1.1 | $5.95 | $1.01 | 1h20m | 3m21s |
| CUA-bench | $188.04 | $280.03 | 14.71s | 36.55s |
| Time Horizon Index: KSP | $1303.37 | $543.14 | N/A | N/A |
Results available only for Claude Opus 5.5
- SRE Bench
Results available only for Gemini 3.8 Flash
- LegalBench
- MortgageTax
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- LiveCodeBench
- SkillsBench
- SWE-bench