Model comparison

Claude Sonnet 5 vs GPT-5.6 Luna: Benchmark Comparison

Claude Sonnet 5 has the higher score on 12 of 27 shared benchmarks; GPT-5.6 Luna leads on 14.

The largest observed score gap is 17.00 pts on ProofBench v1.1 , where Claude Sonnet 5 leads.

Reported ±1 standard-error ranges overlap on 12 of 27 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 GPT-5.6 Luna Gap Reported uncertainty
Vals Index 51.77% ±1.09 51.69% ±1.06 0.09 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 1.25% ±0.83 3.75 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 41.83% ±3.43 36.54% ±3.35 5.29 pts Reported ±1 SE ranges overlap
LegalBench 83.92% ±0.46 84.03% ±0.42 0.11 pts Reported ±1 SE ranges overlap
EMB 66.32% ±3.01 67.12% ±2.94 0.80 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 55.04% ±0.31 1.13 pts Reported ±1 SE ranges do not overlap
MortgageTax 70.03% ±0.90 67.29% ±0.92 2.74 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 60.81% ±3.26 1.46 pts Reported ±1 SE ranges overlap
TaxEval v2 75.63% ±0.84 76.17% ±0.84 0.53 pts Reported ±1 SE ranges overlap
MedCode 47.54% ±2.27 42.39% ±2.27 5.15 pts Reported ±1 SE ranges do not overlap
MedScribe 76.05% ±3.05 84.39% ±2.58 8.34 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 60.00% ±4.92 17.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 2.86% ±2.01 0.00% ±0.00 2.86 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 88.89% ±2.22 91.67% ±1.74 2.78 pts Reported ±1 SE ranges overlap
MMLU Pro 87.55% ±0.37 86.04% ±0.35 1.51 pts Reported ±1 SE ranges do not overlap
MMMU Pro 83.01% ±0.90 85.03% ±0.86 2.02 pts Reported ±1 SE ranges do not overlap
SAGE 48.92% ±3.40 44.22% ±3.33 4.70 pts Reported ±1 SE ranges overlap
Code Migration 44.39% ±4.25 44.55% ±4.24 0.16 pts Reported ±1 SE ranges overlap
IOI 45.00% ±2.75 61.78% ±11.61 16.78 pts Reported ±1 SE ranges do not overlap
ProgramBench 0.00% ±0.00 0.00% ±0.00 0.00 pts Reported ±1 SE ranges overlap
SkillsBench 46.48% ±4.49 60.45% ±4.75 13.97 pts Reported ±1 SE ranges do not overlap
SWE-bench 79.60% ±1.80 93.00% ±1.14 13.40 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 9.60% ±1.01 11.62% ±1.01 2.02 pts Reported ±1 SE ranges overlap
Vibe Code Bench 1-100 13.82% ±2.98 22.59% ±4.00 8.77 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 81.33% ±3.05 77.06% ±3.07 4.27 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 73.63% ±5.56 11.73 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 66.03% ±1.23 61.16% ±1.27 4.87 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Sonnet 5 average GPT-5.6 Luna average
Legal 43.58% 40.61%
Finance 65.63% 65.29%
Healthcare 61.80% 63.39%
Math 77.00% 60.00%
Science 2.86% 0.00%
Academic 86.48% 87.58%
Education 48.92% 44.22%
Coding 40.03% 46.38%
Beta 61.91% 73.63%
Social Mobility 66.03% 61.16%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost GPT-5.6 Luna cost Claude Sonnet 5 latency GPT-5.6 Luna latency
Vals Index $13.72 $0.82 54m16s 29m27s
Harvey's Legal Agent Benchmark $8.95 $0.38 38m52s 10m56s
Legal Research Bench $2.72 $0.85 25m46s 39m34s
LegalBench N/A N/A 4.90s 5.89s
EMB $10.29 $0.37 49m46s 11m00s
Finance Agent (v2) $0.75 $0.28 13m12s 12m52s
MortgageTax N/A N/A 28.67s 22.84s
Tax Agent Bench $1.79 $0.47 19m19s 42m46s
TaxEval v2 N/A N/A 3m22s 76.52s
MedCode N/A N/A 2m15s 81.28s
MedScribe N/A N/A 4m12s 116.89s
ProofBench v1.1 $1.37 $0.12 14m58s 8m34s
Terminal-Bench Science $18.70 $0.58 2h55m 2h03m
GPQA Diamond N/A N/A 63.23s 53.52s
MMLU Pro N/A N/A 25.19s 16.95s
MMMU Pro N/A N/A 18.82s 35.10s
SAGE N/A N/A 7m14s 51.88s
Code Migration $35.31 $1.88 1h57m 58m35s
IOI $12.89 $0.58 59m44s 1h10m
ProgramBench $24.36 $0.94 1h31m 41m34s
SkillsBench $2.90 $0.21 15m12s 5m20s
SWE-bench $1.49 $0.04 16m02s 3m21s
Terminal-Bench 4.0 $26.33 $0.73 1h45m 35m55s
Vibe Code Bench 1-100 $71.15 $4.91 4h16m 2h03m
Vibe Code Bench v1.1 $25.39 $0.73 1h07m 25m52s
CyberBench v1.1 $1.89 $0.40 16m57s 16m15s
Public Benefits Bench v1.1 $1.29 $0.27 21m09s 38m40s

Results available only for Claude Sonnet 5

  • LiveCodeBench

Results available only for GPT-5.6 Luna

  • BioMysteryBench
  • MysteryMechanism
Model details Claude Sonnet 5 Model details GPT-5.6 Luna