Claude Opus 5 vs GPT-5.6 Sol: Benchmark Comparison

Claude Opus 5 has the higher score on 28 of 34 shared benchmarks; GPT-5.6 Sol leads on 6.

The largest observed score gap is 19.60 pts on MedCode , where Claude Opus 5 leads.

Reported ±1 standard-error ranges overlap on 13 of 32 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5 GPT-5.6 Sol Gap Reported uncertainty
Vals Index 63.67% ±0.96 58.01% ±1.02 5.67 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 33.02% 23.88% 9.14 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 6.67% ±1.65 2.50% ±0.83 4.17 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 48.08% ±3.47 7.21 pts Reported ±1 SE ranges do not overlap
LegalBench 86.97% ±0.42 86.97% ±0.41 0.01 pts Reported ±1 SE ranges overlap
EMB 73.56% ±2.24 72.34% ±2.37 1.22 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 58.63% ±0.08 53.76% ±0.85 4.88 pts Reported ±1 SE ranges do not overlap
MortgageTax 72.06% ±0.88 67.29% ±0.92 4.77 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 75.06% ±2.96 67.95% ±3.12 7.11 pts Reported ±1 SE ranges do not overlap
TaxEval v2 75.14% ±0.83 74.78% ±0.86 0.37 pts Reported ±1 SE ranges overlap
MedCode 63.57% ±1.99 43.97% ±2.26 19.60 pts Reported ±1 SE ranges do not overlap
MedScribe 90.98% ±1.92 85.23% ±1.97 5.75 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 99.00% ±1.00 83.00% ±3.77 16.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 79.26% ±2.06 71.11% ±0.64 8.15 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 37.39% ±3.25 33.33% ±3.17 4.05 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 27.14% ±5.35 20.00% ±4.82 7.14 pts Reported ±1 SE ranges overlap
GPQA Diamond 93.43% ±1.24 95.20% ±1.07 1.77 pts Reported ±1 SE ranges overlap
MMLU Pro 91.59% ±0.28 89.10% ±0.31 2.49 pts Reported ±1 SE ranges do not overlap
MMMU Pro 89.88% ±0.72 88.84% ±0.76 1.04 pts Reported ±1 SE ranges overlap
SAGE 49.43% ±3.29 52.56% ±3.42 3.14 pts Reported ±1 SE ranges overlap
Code Migration 57.47% ±4.37 52.92% ±4.35 4.56 pts Reported ±1 SE ranges overlap
IOI 84.33% ±9.96 91.17% ±4.51 6.83 pts Reported ±1 SE ranges overlap
LiveCodeBench 89.03% ±0.91 82.60% ±1.09 6.43 pts Reported ±1 SE ranges do not overlap
ProgramBench 3.00% ±1.21 1.50% ±0.86 1.50 pts Reported ±1 SE ranges overlap
SkillsBench 60.44% ±4.58 54.10% ±4.73 6.34 pts Reported ±1 SE ranges overlap
SWE-bench 97.00% ±0.76 96.20% ±0.86 0.80 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 53.53% ±1.34 37.88% ±0.00 15.66 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 28.53% ±4.33 20.01% ±3.63 8.52 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 88.40% ±3.00 80.50% ±3.72 7.91 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 65.36% ±5.55 76.31% ±5.18 10.95 pts Reported ±1 SE ranges do not overlap
SRE Bench 12.21% ±2.03 30.53% ±2.85 18.32 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 76.93% ±1.10 66.51% ±1.23 10.42 pts Reported ±1 SE ranges do not overlap
CUA-bench 9.00% 8.33% 0.67 pts Uncertainty comparison unavailable
Time Horizon Index: KSP 18.83% ±0.00 23.83% ±0.00 5.00 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Opus 5 average GPT-5.6 Sol average
Legal 49.64% 45.85%
Finance 70.89% 67.22%
Healthcare 77.28% 64.60%
Math 99.00% 83.00%
Science 47.93% 41.48%
Academic 91.64% 91.05%
Education 49.43% 52.56%
Coding 62.42% 57.43%
Cyber 38.79% 53.42%
Social Mobility 76.93% 66.51%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5 cost GPT-5.6 Sol cost Claude Opus 5 latency GPT-5.6 Sol latency
Vals Index $19.31 $14.24 57m41s 37m23s
Vals RSI Index $418.69 $347.10 90h00m 90h00m
Harvey's Legal Agent Benchmark $23.67 $10.37 55m37s 22m50s
Legal Research Bench $6.76 $21.61 31m01s 1h17m
LegalBench N/A N/A 4.42s 6.20s
EMB $6.07 $6.01 25m59s 9m50s
Finance Agent (v2) $5.12 $1.25 9m58s 19m25s
MortgageTax N/A N/A 7.98s 12.62s
Tax Agent Bench $5.14 $7.51 16m48s 44m41s
TaxEval v2 N/A N/A 58.34s 62.03s
MedCode N/A N/A 24.50s 96.59s
MedScribe N/A N/A 76.56s 94.40s
ProofBench v1.1 $2.11 $1.44 9m51s 8m40s
BioMysteryBench $3.29 $1.93 36m36s 18m18s
MysteryMechanism $3.28 $0.96 16m53s 8m43s
Terminal-Bench Science $32.54 $7.32 2h12m 1h44m
GPQA Diamond N/A N/A 39.59s 58.30s
MMLU Pro N/A N/A 9.60s 15.65s
MMMU Pro N/A N/A 30.03s 26.20s
SAGE N/A N/A 56.17s 75.39s
Code Migration $60.51 $24.54 2h51m 59m12s
IOI $16.48 $7.98 56m42s 1h02m
LiveCodeBench N/A N/A 59.22s 56.51s
ProgramBench $60.29 $15.29 2h55m 29m24s
SkillsBench $2.51 $4.14 8m25s 13m07s
SWE-bench $1.29 $1.15 9m37s 3m02s
Terminal-Bench 4.0 $18.60 $7.98 1h07m 33m46s
Vibe Code Bench 1-100 $41.49 $47.43 1h22m 2h37m
Vibe Code Bench v1.1 $33.88 $33.40 1h28m 33m53s
CyberBench v1.1 $2.87 $3.05 19m25s 10m47s
SRE Bench $23.63 $42.52 1h21m 1h28m
Public Benefits Bench v1.1 $3.95 $8.85 30m09s 1h01m
CUA-bench $232.23 $206.99 25.29s 18.80s
Time Horizon Index: KSP $1348.77 $2027.20 N/A N/A
Model details Claude Opus 5 Model details GPT-5.6 Sol