Claude Fable 5.1 vs GPT-5.6 Sol: Benchmark Comparison

Claude Fable 5.1 has the higher score on 27 of 32 shared benchmarks; GPT-5.6 Sol leads on 5.

The largest observed score gap is 39.50 pts on Time Horizon Index: KSP , where Claude Fable 5.1 leads.

Reported ±1 standard-error ranges overlap on 9 of 30 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5.1 GPT-5.6 Sol Gap Reported uncertainty
Vals Index 65.83% ±1.11 58.01% ±1.02 7.82 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 36.09% 23.88% 12.21 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 6.67% ±1.17 2.50% ±0.83 4.17 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 55.29% ±3.46 48.08% ±3.47 7.21 pts Reported ±1 SE ranges do not overlap
LegalBench 88.51% ±0.42 86.97% ±0.41 1.55 pts Reported ±1 SE ranges do not overlap
EMB 76.67% ±2.08 72.34% ±2.37 4.33 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 58.88% ±2.06 53.76% ±0.85 5.12 pts Reported ±1 SE ranges do not overlap
MortgageTax 70.79% ±0.90 67.29% ±0.92 3.50 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 77.64% ±2.83 67.95% ±3.12 9.69 pts Reported ±1 SE ranges do not overlap
TaxEval v2 75.96% ±0.83 74.78% ±0.86 1.19 pts Reported ±1 SE ranges overlap
MedCode 53.51% ±2.17 43.97% ±2.26 9.54 pts Reported ±1 SE ranges do not overlap
MedScribe 91.29% ±1.95 85.23% ±1.97 6.06 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 100.00% ±0.00 83.00% ±3.77 17.00 pts Reported ±1 SE ranges do not overlap
MysteryMechanism 47.75% ±3.36 33.33% ±3.17 14.41 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 40.00% ±5.90 20.00% ±4.82 20.00 pts Reported ±1 SE ranges do not overlap
GPQA Diamond 93.43% ±1.86 95.20% ±1.07 1.77 pts Reported ±1 SE ranges overlap
MMLU Pro 92.38% ±0.27 89.10% ±0.31 3.28 pts Reported ±1 SE ranges do not overlap
MMMU Pro 90.64% ±0.70 88.84% ±0.76 1.79 pts Reported ±1 SE ranges do not overlap
SAGE 48.53% ±3.34 52.56% ±3.42 4.04 pts Reported ±1 SE ranges overlap
Code Migration 54.61% ±4.81 52.92% ±4.35 1.69 pts Reported ±1 SE ranges overlap
IOI 90.78% ±4.65 91.17% ±4.51 0.39 pts Reported ±1 SE ranges overlap
LiveCodeBench 90.52% ±0.86 82.60% ±1.09 7.91 pts Reported ±1 SE ranges do not overlap
ProgramBench 7.00% ±1.81 1.50% ±0.86 5.50 pts Reported ±1 SE ranges do not overlap
SkillsBench 61.55% ±4.76 54.10% ±4.73 7.45 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 58.08% ±3.31 37.88% ±0.00 20.20 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench 1-100 28.00% ±4.49 20.01% ±3.63 7.99 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 90.26% ±1.57 80.50% ±3.72 9.77 pts Reported ±1 SE ranges do not overlap
CyberBench v1.1 70.42% ±5.43 76.31% ±5.18 5.89 pts Reported ±1 SE ranges overlap
SRE Bench 22.90% ±2.60 30.53% ±2.85 7.63 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 74.90% ±1.13 66.51% ±1.23 8.39 pts Reported ±1 SE ranges do not overlap
CUA-bench 13.17% 8.33% 4.83 pts Uncertainty comparison unavailable
Time Horizon Index: KSP 63.33% ±0.00 23.83% ±0.00 39.50 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Fable 5.1 average GPT-5.6 Sol average
Legal 50.16% 45.85%
Finance 71.99% 67.22%
Healthcare 72.40% 64.60%
Math 100.00% 83.00%
Science 43.87% 26.67%
Academic 92.15% 91.05%
Education 48.53% 52.56%
Coding 60.10% 52.58%
Cyber 46.66% 53.42%
Social Mobility 74.90% 66.51%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5.1 cost GPT-5.6 Sol cost Claude Fable 5.1 latency GPT-5.6 Sol latency
Vals Index $28.71 $14.24 1h17m 37m23s
Vals RSI Index $298.22 $347.10 90h00m 90h00m
Harvey's Legal Agent Benchmark $46.21 $10.37 1h52m 22m50s
Legal Research Bench $23.06 $21.61 57m29s 1h17m
LegalBench N/A N/A 12.40s 6.20s
EMB $15.93 $6.01 29m32s 9m50s
Finance Agent (v2) $8.35 $1.25 18m14s 19m25s
MortgageTax N/A N/A 14.98s 12.62s
Tax Agent Bench $13.18 $7.51 28m34s 44m41s
TaxEval v2 N/A N/A 85.56s 62.03s
MedCode N/A N/A 3m35s 96.59s
MedScribe N/A N/A 3m08s 94.40s
ProofBench v1.1 $2.70 $1.44 7m58s 8m40s
MysteryMechanism $5.63 $0.96 17m41s 8m43s
Terminal-Bench Science $38.01 $7.32 2h36m 1h44m
GPQA Diamond N/A N/A 45.86s 58.30s
MMLU Pro N/A N/A 20.59s 15.65s
MMMU Pro N/A N/A 37.47s 26.20s
SAGE N/A N/A 105.75s 75.39s
Code Migration $70.97 $24.54 4h20m 59m12s
IOI $11.20 $7.98 30m06s 1h02m
LiveCodeBench N/A N/A 59.74s 56.51s
ProgramBench $58.52 $15.29 2h27m 29m24s
SkillsBench $4.52 $4.14 11m02s 13m07s
Terminal-Bench 4.0 $17.18 $7.98 1h00m 33m46s
Vibe Code Bench 1-100 $149.66 $47.43 2h53m 2h37m
Vibe Code Bench v1.1 $33.37 $33.40 57m40s 33m53s
CyberBench v1.1 $4.01 $3.05 18m32s 10m47s
SRE Bench $32.23 $42.52 2h03m 1h28m
Public Benefits Bench v1.1 $7.48 $8.85 48m02s 1h01m
CUA-bench $323.78 $206.99 37.44s 18.80s
Time Horizon Index: KSP $4655.01 $2027.20 N/A N/A

Results available only for Claude Fable 5.1

None.

Results available only for GPT-5.6 Sol

  • BioMysteryBench
  • SWE-bench
Model details Claude Fable 5.1 Model details GPT-5.6 Sol