Claude Fable 5 vs GPT-5.6 Sol: Benchmark Comparison

Claude Fable 5 has the higher score on 21 of 25 shared benchmarks; GPT-5.6 Sol leads on 4.

The largest observed score gap is 12.10 pts on MedCode , where Claude Fable 5 leads.

Reported ±1 standard-error ranges overlap on 12 of 24 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5 GPT-5.6 Sol Gap Reported uncertainty
Vals Index 61.39% ±1.00 58.01% ±1.02 3.38 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 24.32% 23.88% 0.44 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 11.25% ±2.15 2.50% ±0.83 8.75 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 49.52% ±3.48 48.08% ±3.47 1.44 pts Reported ±1 SE ranges overlap
LegalBench 88.56% ±0.33 86.97% ±0.41 1.60 pts Reported ±1 SE ranges do not overlap
EMB 73.67% ±2.47 72.34% ±2.37 1.33 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 56.31% ±0.84 53.76% ±0.85 2.56 pts Reported ±1 SE ranges do not overlap
MortgageTax 68.92% ±0.91 67.29% ±0.92 1.63 pts Reported ±1 SE ranges overlap
Tax Agent Bench 69.82% ±3.12 67.95% ±3.12 1.87 pts Reported ±1 SE ranges overlap
TaxEval v2 76.94% ±0.82 74.78% ±0.86 2.17 pts Reported ±1 SE ranges do not overlap
MedCode 56.07% ±2.20 43.97% ±2.26 12.10 pts Reported ±1 SE ranges do not overlap
MedScribe 88.52% ±1.95 85.23% ±1.97 3.29 pts Reported ±1 SE ranges overlap
ProofBench v1.1 95.00% ±2.19 83.00% ±3.77 12.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 15.71% ±4.38 20.00% ±4.82 4.29 pts Reported ±1 SE ranges overlap
GPQA Diamond 93.18% ±1.94 95.20% ±1.07 2.02 pts Reported ±1 SE ranges overlap
MMLU Pro 91.50% ±0.28 89.10% ±0.31 2.40 pts Reported ±1 SE ranges do not overlap
MMMU Pro 89.31% ±0.74 88.84% ±0.76 0.46 pts Reported ±1 SE ranges overlap
SAGE 51.89% ±3.40 52.56% ±3.42 0.67 pts Reported ±1 SE ranges overlap
Code Migration 55.06% ±4.61 52.92% ±4.35 2.15 pts Reported ±1 SE ranges overlap
LiveCodeBench 89.78% ±0.89 82.60% ±1.09 7.17 pts Reported ±1 SE ranges do not overlap
ProgramBench 2.00% ±0.99 1.50% ±0.86 0.50 pts Reported ±1 SE ranges overlap
SWE-bench 95.00% ±0.98 96.20% ±0.86 1.20 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 41.41% ±1.34 37.88% ±0.00 3.54 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 90.35% ±2.10 80.50% ±3.72 9.86 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 70.43% ±1.19 66.51% ±1.23 3.92 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Fable 5 average GPT-5.6 Sol average
Legal 49.78% 45.85%
Finance 69.13% 67.22%
Healthcare 72.30% 64.60%
Math 95.00% 83.00%
Science 15.71% 20.00%
Academic 91.33% 91.05%
Education 51.89% 52.56%
Coding 62.27% 58.60%
Social Mobility 70.43% 66.51%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5 cost GPT-5.6 Sol cost Claude Fable 5 latency GPT-5.6 Sol latency
Vals Index $29.59 $14.24 42m24s 37m23s
Vals RSI Index $491.22 $347.10 90h00m 90h00m
Harvey's Legal Agent Benchmark $19.23 $10.37 26m53s 22m50s
Legal Research Bench $9.79 $21.61 22m18s 1h17m
LegalBench N/A N/A 8.96s 6.20s
EMB $12.35 $6.01 26m34s 9m50s
Finance Agent (v2) $8.06 $1.25 10m11s 19m25s
MortgageTax N/A N/A 16.10s 12.62s
Tax Agent Bench $6.70 $7.51 14m27s 44m41s
TaxEval v2 N/A N/A 56.87s 62.03s
MedCode N/A N/A 91.44s 96.59s
MedScribe N/A N/A 119.47s 94.40s
ProofBench v1.1 $5.45 $1.44 13m43s 8m40s
Terminal-Bench Science $51.99 $7.32 1h59m 1h44m
GPQA Diamond N/A N/A 99.90s 58.30s
MMLU Pro N/A N/A 25.00s 15.65s
MMMU Pro N/A N/A 61.44s 26.20s
SAGE N/A N/A 116.92s 75.39s
Code Migration $112.10 $24.54 1h49m 59m12s
LiveCodeBench N/A N/A 118.53s 56.51s
ProgramBench $75.68 $15.29 2h37m 29m24s
SWE-bench $2.05 $1.15 5m56s 3m02s
Terminal-Bench 4.0 $30.34 $7.98 1h08m 33m46s
Vibe Code Bench v1.1 $41.71 $33.40 1h02m 33m53s
Public Benefits Bench v1.1 $4.59 $8.85 21m17s 1h01m

Results available only for Claude Fable 5

None.

Results available only for GPT-5.6 Sol

  • BioMysteryBench
  • MysteryMechanism
  • IOI
  • SkillsBench
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • SRE Bench
  • CUA-bench
  • Time Horizon Index: KSP
Model details Claude Fable 5 Model details GPT-5.6 Sol