Model comparison

Claude Sonnet 5 vs GPT-6 Luna: Benchmark Comparison

Claude Sonnet 5 has the higher score on 10 of 18 shared benchmarks; GPT-6 Luna leads on 8.

The largest observed score gap is 14.34 pts on CyberBench v1.1 , where GPT-6 Luna leads.

Reported ±1 standard-error ranges overlap on 11 of 18 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Sonnet 5 GPT-6 Luna Gap Reported uncertainty
Vals Index 51.77% ±1.09 51.22% ±1.07 0.56 pts Reported ±1 SE ranges overlap
Harvey's Legal Agent Benchmark 5.00% ±1.65 2.92% ±1.17 2.08 pts Reported ±1 SE ranges overlap
Legal Research Bench 41.83% ±3.43 30.29% ±3.19 11.54 pts Reported ±1 SE ranges do not overlap
EMB 66.32% ±3.01 68.52% ±2.81 2.20 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 53.91% ±0.52 49.87% ±0.23 4.04 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 62.27% ±3.19 58.86% ±3.21 3.41 pts Reported ±1 SE ranges overlap
MedCode 47.54% ±2.27 44.69% ±2.30 2.85 pts Reported ±1 SE ranges overlap
MedScribe 76.05% ±3.05 83.71% ±1.95 7.66 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 77.00% ±4.23 64.00% ±4.82 13.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 2.86% ±2.01 4.29% ±2.44 1.43 pts Reported ±1 SE ranges overlap
SAGE 48.92% ±3.40 48.09% ±3.38 0.83 pts Reported ±1 SE ranges overlap
Code Migration 44.39% ±4.25 42.55% ±4.42 1.84 pts Reported ±1 SE ranges overlap
IOI 45.00% ±2.75 55.56% ±8.87 10.56 pts Reported ±1 SE ranges overlap
ProgramBench 0.00% ±0.00 0.50% ±0.50 0.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 9.60% ±1.01 13.64% ±1.51 4.04 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 81.33% ±3.05 81.65% ±3.38 0.32 pts Reported ±1 SE ranges overlap
CyberBench v1.1 61.91% ±5.74 76.25% ±5.29 14.34 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 66.03% ±1.23 57.65% ±1.28 8.39 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Sonnet 5 average GPT-6 Luna average
Legal 23.41% 16.60%
Finance 60.83% 59.08%
Healthcare 61.80% 64.20%
Math 77.00% 64.00%
Science 2.86% 4.29%
Education 48.92% 48.09%
Coding 36.06% 38.78%
Beta 61.91% 76.25%
Social Mobility 66.03% 57.65%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Sonnet 5 cost GPT-6 Luna cost Claude Sonnet 5 latency GPT-6 Luna latency
Vals Index $13.72 $0.43 54m16s 30m34s
Harvey's Legal Agent Benchmark $8.95 $0.30 38m52s 17m32s
Legal Research Bench $2.72 $0.44 25m46s 42m28s
EMB $10.29 $0.14 49m46s 17m21s
Finance Agent (v2) $0.75 $0.12 13m12s 14m58s
Tax Agent Bench $1.79 $0.19 19m19s 33m16s
MedCode N/A N/A 2m15s 108.77s
MedScribe N/A N/A 4m12s 2m50s
ProofBench v1.1 $1.37 $0.04 14m58s 7m31s
Terminal-Bench Science $18.70 $0.24 2h55m 1h16m
SAGE N/A N/A 7m14s 91.07s
Code Migration $35.31 $0.60 1h57m 50m23s
IOI $12.89 $0.20 59m44s 35m47s
ProgramBench $24.36 $0.18 1h31m 28m55s
Terminal-Bench 4.0 $26.33 $0.35 1h45m 32m46s
Vibe Code Bench v1.1 $25.39 $1.35 1h07m 37m00s
CyberBench v1.1 $1.89 $0.13 16m57s 15m50s
Public Benefits Bench v1.1 $1.29 $0.22 21m09s 52m59s

Results available only for Claude Sonnet 5

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • SWE-bench
  • Vibe Code Bench 1-100

Results available only for GPT-6 Luna

  • BioMysteryBench
  • MysteryMechanism
Model details Claude Sonnet 5 Model details GPT-6 Luna