Claude Fable 5 vs GPT-5.6 Luna: Benchmark Comparison

Claude Fable 5 has the higher score on 23 of 23 shared benchmarks; GPT-5.6 Luna leads on 0.

The largest observed score gap is 35.00 pts on ProofBench v1.1 , where Claude Fable 5 leads.

Reported ±1 standard-error ranges overlap on 5 of 23 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5 GPT-5.6 Luna Gap Reported uncertainty
Vals Index 61.39% ±1.00 51.69% ±1.06 9.70 pts Reported ±1 SE ranges do not overlap
Harvey's Legal Agent Benchmark 11.25% ±2.15 1.25% ±0.83 10.00 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 49.52% ±3.48 36.54% ±3.35 12.98 pts Reported ±1 SE ranges do not overlap
LegalBench 88.56% ±0.33 84.03% ±0.42 4.53 pts Reported ±1 SE ranges do not overlap
EMB 73.67% ±2.47 67.12% ±2.94 6.55 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 56.31% ±0.84 55.04% ±0.31 1.27 pts Reported ±1 SE ranges do not overlap
MortgageTax 68.92% ±0.91 67.29% ±0.92 1.63 pts Reported ±1 SE ranges overlap
Tax Agent Bench 69.82% ±3.12 60.81% ±3.26 9.01 pts Reported ±1 SE ranges do not overlap
TaxEval v2 76.94% ±0.82 76.17% ±0.84 0.78 pts Reported ±1 SE ranges overlap
MedCode 56.07% ±2.20 42.39% ±2.27 13.68 pts Reported ±1 SE ranges do not overlap
MedScribe 88.52% ±1.95 84.39% ±2.58 4.13 pts Reported ±1 SE ranges overlap
ProofBench v1.1 95.00% ±2.19 60.00% ±4.92 35.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 15.71% ±4.38 0.00% ±0.00 15.71 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 93.18% ±1.94 91.67% ±1.74 1.52 pts Reported ±1 SE ranges overlap
MMLU Pro 91.50% ±0.28 86.04% ±0.35 5.47 pts Reported ±1 SE ranges do not overlap
MMMU Pro 89.31% ±0.74 85.03% ±0.86 4.28 pts Reported ±1 SE ranges do not overlap
SAGE 51.89% ±3.40 44.22% ±3.33 7.67 pts Reported ±1 SE ranges do not overlap
Code Migration 55.06% ±4.61 44.55% ±4.24 10.52 pts Reported ±1 SE ranges do not overlap
ProgramBench 2.00% ±0.99 0.00% ±0.00 2.00 pts Reported ±1 SE ranges do not overlap
SWE-bench 95.00% ±0.98 93.00% ±1.14 2.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 41.41% ±1.34 11.62% ±1.01 29.80 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 90.35% ±2.10 77.06% ±3.07 13.30 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 70.43% ±1.19 61.16% ±1.27 9.27 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Fable 5 average GPT-5.6 Luna average
Legal 49.78% 40.61%
Finance 69.13% 65.29%
Healthcare 72.30% 63.39%
Math 95.00% 60.00%
Science 15.71% 0.00%
Academic 91.33% 87.58%
Education 51.89% 44.22%
Coding 56.77% 45.24%
Social Mobility 70.43% 61.16%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5 cost GPT-5.6 Luna cost Claude Fable 5 latency GPT-5.6 Luna latency
Vals Index $29.59 $0.82 42m24s 29m27s
Harvey's Legal Agent Benchmark $19.23 $0.38 26m53s 10m56s
Legal Research Bench $9.79 $0.85 22m18s 39m34s
LegalBench N/A N/A 8.96s 5.89s
EMB $12.35 $0.37 26m34s 11m00s
Finance Agent (v2) $8.06 $0.28 10m11s 12m52s
MortgageTax N/A N/A 16.10s 22.84s
Tax Agent Bench $6.70 $0.47 14m27s 42m46s
TaxEval v2 N/A N/A 56.87s 76.52s
MedCode N/A N/A 91.44s 81.28s
MedScribe N/A N/A 119.47s 116.89s
ProofBench v1.1 $5.45 $0.12 13m43s 8m34s
Terminal-Bench Science $51.99 $0.58 1h59m 2h03m
GPQA Diamond N/A N/A 99.90s 53.52s
MMLU Pro N/A N/A 25.00s 16.95s
MMMU Pro N/A N/A 61.44s 35.10s
SAGE N/A N/A 116.92s 51.88s
Code Migration $112.10 $1.88 1h49m 58m35s
ProgramBench $75.68 $0.94 2h37m 41m34s
SWE-bench $2.05 $0.04 5m56s 3m21s
Terminal-Bench 4.0 $30.34 $0.73 1h08m 35m55s
Vibe Code Bench v1.1 $41.71 $0.73 1h02m 25m52s
Public Benefits Bench v1.1 $4.59 $0.27 21m17s 38m40s

Results available only for Claude Fable 5

  • Vals RSI Index
  • LiveCodeBench

Results available only for GPT-5.6 Luna

  • BioMysteryBench
  • MysteryMechanism
  • IOI
  • SkillsBench
  • Vibe Code Bench 1-100
  • CyberBench v1.1
Model details Claude Fable 5 Model details GPT-5.6 Luna