Claude Fable 5 vs GPT-6 Sol: Benchmark Comparison

Claude Fable 5 has the higher score on 12 of 17 shared benchmarks; GPT-6 Sol leads on 4.

The largest observed score gap is 20.67 pts on Legal Research Bench , where Claude Fable 5 leads.

Reported ±1 standard-error ranges overlap on 5 of 16 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5 GPT-6 Sol Gap Reported uncertainty
Vals Index 61.39% ±1.00 57.54% ±1.01 3.85 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 24.32% 28.11% 3.79 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 11.25% ±2.15 1.67% ±0.82 9.58 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 49.52% ±3.48 28.85% ±3.15 20.67 pts Reported ±1 SE ranges do not overlap
EMB 73.67% ±2.47 71.53% ±2.35 2.14 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 56.31% ±0.84 49.05% ±0.58 7.26 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 69.82% ±3.12 53.05% ±3.29 16.78 pts Reported ±1 SE ranges do not overlap
MedCode 56.07% ±2.20 47.07% ±2.12 9.00 pts Reported ±1 SE ranges do not overlap
MedScribe 88.52% ±1.95 82.03% ±1.94 6.49 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 95.00% ±2.19 83.00% ±3.77 12.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 15.71% ±4.38 30.00% ±5.52 14.29 pts Reported ±1 SE ranges do not overlap
SAGE 51.89% ±3.40 44.79% ±3.36 7.10 pts Reported ±1 SE ranges do not overlap
Code Migration 55.06% ±4.61 57.20% ±4.21 2.13 pts Reported ±1 SE ranges overlap
ProgramBench 2.00% ±0.99 2.00% ±0.99 0.00 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 41.41% ±1.34 44.44% ±3.54 3.03 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 90.35% ±2.10 87.82% ±2.53 2.53 pts Reported ±1 SE ranges overlap
Public Benefits Bench v1.1 70.43% ±1.19 56.63% ±1.29 13.80 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Fable 5 average GPT-6 Sol average
Legal 30.38% 15.26%
Finance 66.60% 57.88%
Healthcare 72.30% 64.55%
Math 95.00% 83.00%
Science 15.71% 30.00%
Education 51.89% 44.79%
Coding 47.21% 47.87%
Social Mobility 70.43% 56.63%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5 cost GPT-6 Sol cost Claude Fable 5 latency GPT-6 Sol latency
Vals Index $29.59 $7.58 42m24s 29m11s
Vals RSI Index $491.22 $665.11 90h00m 90h00m
Harvey's Legal Agent Benchmark $19.23 $3.36 26m53s 13m40s
Legal Research Bench $9.79 $4.90 22m18s 24m09s
EMB $12.35 $1.28 26m34s 10m22s
Finance Agent (v2) $8.06 $2.12 10m11s 9m00s
Tax Agent Bench $6.70 $2.28 14m27s 17m32s
MedCode N/A N/A 91.44s 62.71s
MedScribe N/A N/A 119.47s 81.14s
ProofBench v1.1 $5.45 $0.48 13m43s 5m05s
Terminal-Bench Science $51.99 $5.82 1h59m 1h22m
SAGE N/A N/A 116.92s 31.42s
Code Migration $112.10 $15.68 1h49m 1h26m
ProgramBench $75.68 $11.67 2h37m 46m11s
Terminal-Bench 4.0 $30.34 $5.79 1h08m 34m56s
Vibe Code Bench v1.1 $41.71 $26.36 1h02m 38m56s
Public Benefits Bench v1.1 $4.59 $2.49 21m17s 31m02s

Results available only for Claude Fable 5

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SWE-bench

Results available only for GPT-6 Sol

  • BioMysteryBench
  • MysteryMechanism
  • IOI
  • CyberBench v1.1
Model details Claude Fable 5 Model details GPT-6 Sol