Claude Opus 5 vs GPT-6 Sol: Benchmark Comparison
Claude Opus 5 has the higher score on 19 of 21 shared benchmarks; GPT-6 Sol leads on 2.
The largest observed score gap is 26.44 pts on Legal Research Bench , where Claude Opus 5 leads.
Reported ±1 standard-error ranges overlap on 8 of 20 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Claude Opus 5 | GPT-6 Sol | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 63.67% ±0.96 | 57.54% ±1.01 | 6.14 pts | Reported ±1 SE ranges do not overlap |
| Vals RSI Index | 33.02% | 28.11% | 4.91 pts | Uncertainty comparison unavailable |
| Harvey's Legal Agent Benchmark | 6.67% ±1.65 | 1.67% ±0.82 | 5.00 pts | Reported ±1 SE ranges do not overlap |
| Legal Research Bench | 55.29% ±3.46 | 28.85% ±3.15 | 26.44 pts | Reported ±1 SE ranges do not overlap |
| EMB | 73.56% ±2.24 | 71.53% ±2.35 | 2.03 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 58.63% ±0.08 | 49.05% ±0.58 | 9.58 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 75.06% ±2.96 | 53.05% ±3.29 | 22.02 pts | Reported ±1 SE ranges do not overlap |
| MedCode | 63.57% ±1.99 | 47.07% ±2.12 | 16.50 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 90.98% ±1.92 | 82.03% ±1.94 | 8.95 pts | Reported ±1 SE ranges do not overlap |
| ProofBench v1.1 | 99.00% ±1.00 | 83.00% ±3.77 | 16.00 pts | Reported ±1 SE ranges do not overlap |
| BioMysteryBench | 79.26% ±2.06 | 74.81% ±2.43 | 4.44 pts | Reported ±1 SE ranges overlap |
| MysteryMechanism | 37.39% ±3.25 | 30.18% ±3.09 | 7.21 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 27.14% ±5.35 | 30.00% ±5.52 | 2.86 pts | Reported ±1 SE ranges overlap |
| SAGE | 49.43% ±3.29 | 44.79% ±3.36 | 4.63 pts | Reported ±1 SE ranges overlap |
| Code Migration | 57.47% ±4.37 | 57.20% ±4.21 | 0.28 pts | Reported ±1 SE ranges overlap |
| IOI | 84.33% ±9.96 | 82.61% ±9.24 | 1.72 pts | Reported ±1 SE ranges overlap |
| ProgramBench | 3.00% ±1.21 | 2.00% ±0.99 | 1.00 pts | Reported ±1 SE ranges overlap |
| Terminal-Bench 4.0 | 53.53% ±1.34 | 44.44% ±3.54 | 9.09 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 88.40% ±3.00 | 87.82% ±2.53 | 0.58 pts | Reported ±1 SE ranges overlap |
| CyberBench v1.1 | 65.36% ±5.55 | 77.98% ±5.11 | 12.62 pts | Reported ±1 SE ranges do not overlap |
| Public Benefits Bench v1.1 | 76.93% ±1.10 | 56.63% ±1.29 | 20.30 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Claude Opus 5 average | GPT-6 Sol average |
|---|---|---|
| Legal | 30.98% | 15.26% |
| Finance | 69.08% | 57.88% |
| Healthcare | 77.28% | 64.55% |
| Math | 99.00% | 83.00% |
| Science | 47.93% | 45.00% |
| Education | 49.43% | 44.79% |
| Coding | 57.35% | 54.81% |
| Cyber | 65.36% | 77.98% |
| Social Mobility | 76.93% | 56.63% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Claude Opus 5 cost | GPT-6 Sol cost | Claude Opus 5 latency | GPT-6 Sol latency |
|---|---|---|---|---|
| Vals Index | $19.31 | $7.58 | 57m41s | 29m11s |
| Vals RSI Index | $418.69 | $665.11 | 90h00m | 90h00m |
| Harvey's Legal Agent Benchmark | $23.67 | $3.36 | 55m37s | 13m40s |
| Legal Research Bench | $6.76 | $4.90 | 31m01s | 24m09s |
| EMB | $6.07 | $1.28 | 25m59s | 10m22s |
| Finance Agent (v2) | $5.12 | $2.12 | 9m58s | 9m00s |
| Tax Agent Bench | $5.14 | $2.28 | 16m48s | 17m32s |
| MedCode | N/A | N/A | 24.50s | 62.71s |
| MedScribe | N/A | N/A | 76.56s | 81.14s |
| ProofBench v1.1 | $2.11 | $0.48 | 9m51s | 5m05s |
| BioMysteryBench | $3.29 | $0.66 | 36m36s | 5m24s |
| MysteryMechanism | $3.28 | $0.46 | 16m53s | 6m13s |
| Terminal-Bench Science | $32.54 | $5.82 | 2h12m | 1h22m |
| SAGE | N/A | N/A | 56.17s | 31.42s |
| Code Migration | $60.51 | $15.68 | 2h51m | 1h26m |
| IOI | $16.48 | $2.70 | 56m42s | 30m38s |
| ProgramBench | $60.29 | $11.67 | 2h55m | 46m11s |
| Terminal-Bench 4.0 | $18.60 | $5.79 | 1h07m | 34m56s |
| Vibe Code Bench v1.1 | $33.88 | $26.36 | 1h28m | 38m56s |
| CyberBench v1.1 | $2.87 | $1.34 | 19m25s | 12m14s |
| Public Benefits Bench v1.1 | $3.95 | $2.49 | 30m09s | 31m02s |
Results available only for Claude Opus 5
- LegalBench
- MortgageTax
- TaxEval v2
- GPQA Diamond
- MMLU Pro
- MMMU Pro
- LiveCodeBench
- SkillsBench
- SWE-bench
- Vibe Code Bench 1-100
- SRE Bench
- CUA-bench
- Time Horizon Index: KSP
Results available only for GPT-6 Sol
None.