Claude Opus 5 vs GPT-5.6 Luna: Benchmark Comparison

Claude Opus 5 has the higher score on 26 of 29 shared benchmarks; GPT-5.6 Luna leads on 3.

The largest observed score gap is 41.92 pts on Terminal-Bench 4.0 , where Claude Opus 5 leads.

Reported ±1 standard-error ranges overlap on 6 of 29 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Opus 5 GPT-5.6 Luna Gap Reported uncertainty
Vals Index 63.67% ±0.96 51.69% ±1.06 11.99 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 6.67% ±1.65 1.25% ±0.83 5.42 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 36.54% ±3.35 18.75 pts Reported ±1 SE ranges do not overlap
LegalBench 86.97% ±0.42 84.03% ±0.42 2.94 pts Reported ±1 SE ranges do not overlap
EMB 73.56% ±2.24 67.12% ±2.94 6.44 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.63% ±0.08 55.04% ±0.31 3.59 pts Reported ±1 SE ranges do not overlap
MortgageTax 72.06% ±0.88 67.29% ±0.92 4.77 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 75.06% ±2.96 60.81% ±3.26 14.25 pts Reported ±1 SE ranges do not overlap
TaxEval v2 75.14% ±0.83 76.17% ±0.84 1.02 pts Reported ±1 SE ranges overlap
MedCode 63.57% ±1.99 42.39% ±2.27 21.18 pts Reported ±1 SE ranges do not overlap
MedScribe 90.98% ±1.92 84.39% ±2.58 6.59 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 99.00% ±1.00 60.00% ±4.92 39.00 pts Reported ±1 SE ranges do not overlap
BioMysteryBench 79.26% ±2.06 61.48% ±0.74 17.78 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 37.39% ±3.25 14.41% ±2.36 22.97 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 27.14% ±5.35 0.00% ±0.00 27.14 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 93.43% ±1.24 91.67% ±1.74 1.77 pts Reported ±1 SE ranges overlap
MMLU Pro 91.59% ±0.28 86.04% ±0.35 5.55 pts Reported ±1 SE ranges do not overlap
MMMU Pro 89.88% ±0.72 85.03% ±0.86 4.86 pts Reported ±1 SE ranges do not overlap
SAGE 49.43% ±3.29 44.22% ±3.33 5.21 pts Reported ±1 SE ranges overlap
Code Migration 57.47% ±4.37 44.55% ±4.24 12.92 pts Reported ±1 SE ranges do not overlap
IOI 84.33% ±9.96 61.78% ±11.61 22.55 pts Reported ±1 SE ranges do not overlap
ProgramBench 3.00% ±1.21 0.00% ±0.00 3.00 pts Reported ±1 SE ranges do not overlap
SkillsBench 60.44% ±4.58 60.45% ±4.75 0.01 pts Reported ±1 SE ranges overlap
SWE-bench 97.00% ±0.76 93.00% ±1.14 4.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 53.53% ±1.34 11.62% ±1.01 41.92 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 28.53% ±4.33 22.59% ±4.00 5.94 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 88.40% ±3.00 77.06% ±3.07 11.35 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 65.36% ±5.55 73.63% ±5.56 8.27 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 76.93% ±1.10 61.16% ±1.27 15.76 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Opus 5 average GPT-5.6 Luna average
Legal 49.64% 40.61%
Finance 70.89% 65.29%
Healthcare 77.28% 63.39%
Math 99.00% 60.00%
Science 47.93% 25.30%
Academic 91.64% 87.58%
Education 49.43% 44.22%
Coding 59.09% 46.38%
Cyber 65.36% 73.63%
Social Mobility 76.93% 61.16%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Opus 5 cost GPT-5.6 Luna cost Claude Opus 5 latency GPT-5.6 Luna latency
Vals Index $19.31 $0.82 57m41s 29m27s
Harvey's Legal Agent Benchmark $23.67 $0.38 55m37s 10m56s
Legal Research Bench $6.76 $0.85 31m01s 39m34s
LegalBench N/A N/A 4.42s 5.89s
EMB $6.07 $0.37 25m59s 11m00s
Finance Agent (v2) $5.12 $0.28 9m58s 12m52s
MortgageTax N/A N/A 7.98s 22.84s
Tax Agent Bench $5.14 $0.47 16m48s 42m46s
TaxEval v2 N/A N/A 58.34s 76.52s
MedCode N/A N/A 24.50s 81.28s
MedScribe N/A N/A 76.56s 116.89s
ProofBench v1.1 $2.11 $0.12 9m51s 8m34s
BioMysteryBench $3.29 $0.09 36m36s 18m31s
MysteryMechanism $3.28 $0.11 16m53s 8m13s
Terminal-Bench Science $32.54 $0.58 2h12m 2h03m
GPQA Diamond N/A N/A 39.59s 53.52s
MMLU Pro N/A N/A 9.60s 16.95s
MMMU Pro N/A N/A 30.03s 35.10s
SAGE N/A N/A 56.17s 51.88s
Code Migration $60.51 $1.88 2h51m 58m35s
IOI $16.48 $0.58 56m42s 1h10m
ProgramBench $60.29 $0.94 2h55m 41m34s
SkillsBench $2.51 $0.21 8m25s 5m20s
SWE-bench $1.29 $0.04 9m37s 3m21s
Terminal-Bench 4.0 $18.60 $0.73 1h07m 35m55s
Vibe Code Bench 1-100 $41.49 $4.91 1h22m 2h03m
Vibe Code Bench v1.1 $33.88 $0.73 1h28m 25m52s
CyberBench v1.1 $2.87 $0.40 19m25s 16m15s
Public Benefits Bench v1.1 $3.95 $0.27 30m09s 38m40s

Results available only for Claude Opus 5

  • Vals RSI Index
  • LiveCodeBench
  • SRE Bench
  • CUA-bench
  • Time Horizon Index: KSP

Results available only for GPT-5.6 Luna

None.

Model details Claude Opus 5 Model details GPT-5.6 Luna