Model comparison

Claude Fable 5.1 vs GPT-6 Astra: Benchmark Comparison

Claude Fable 5.1 has the higher score on 15 of 23 shared benchmarks; GPT-6 Astra leads on 8.

The largest observed score gap is 33.97 pts on SRE Bench , where GPT-6 Astra leads.

Reported ±1 standard-error ranges overlap on 9 of 21 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5.1 GPT-6 Astra Gap Reported uncertainty
Vals Index 65.83% ±1.11 63.13% ±1.19 2.70 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 36.09% 27.07% 9.02 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 6.67% ±1.17 5.42% ±1.17 1.25 pts Reported ±1 SE ranges overlap
Legal Research Bench 55.29% ±3.46 39.42% ±3.40 15.86 pts Reported ±1 SE ranges do not overlap
EMB 76.67% ±2.08 71.70% ±2.48 4.97 pts Reported ±1 SE ranges do not overlap
Finance Agent (v2) 58.88% ±2.06 53.54% ±2.08 5.34 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 77.64% ±2.83 63.34% ±3.15 14.31 pts Reported ±1 SE ranges do not overlap
MedCode 53.51% ±2.17 48.49% ±2.13 5.02 pts Reported ±1 SE ranges do not overlap
MedScribe 91.29% ±1.95 87.91% ±1.94 3.39 pts Reported ±1 SE ranges overlap
ProofBench v1.1 100.00% ±0.00 99.00% ±1.00 1.00 pts Reported ±1 SE ranges overlap
MysteryMechanism 47.75% ±3.36 53.15% ±3.36 5.41 pts Reported ±1 SE ranges overlap
Terminal-Bench Science 40.00% ±5.90 62.86% ±5.82 22.86 pts Reported ±1 SE ranges do not overlap
SAGE 48.53% ±3.34 46.37% ±3.44 2.16 pts Reported ±1 SE ranges overlap
Code Migration 54.61% ±4.81 67.74% ±4.22 13.13 pts Reported ±1 SE ranges do not overlap
IOI 90.78% ±4.65 100.00% ±0.00 9.22 pts Reported ±1 SE ranges do not overlap
ProgramBench 7.00% ±1.81 5.50% ±1.62 1.50 pts Reported ±1 SE ranges overlap
Terminal-Bench 4.0 58.08% ±3.31 59.60% ±4.40 1.51 pts Reported ±1 SE ranges overlap
Vibe Code Bench 1-100 28.00% ±4.49 27.64% ±4.08 0.36 pts Reported ±1 SE ranges overlap
Vibe Code Bench v1.1 90.26% ±1.57 89.59% ±2.17 0.67 pts Reported ±1 SE ranges overlap
CyberBench v1.1 70.42% ±5.43 41.07% ±2.56 29.34 pts Reported ±1 SE ranges do not overlap
SRE Bench 22.90% ±2.60 56.87% ±3.07 33.97 pts Reported ±1 SE ranges do not overlap
CUA-bench 13.17% 19.17% 6.00 pts Uncertainty comparison unavailable
Time Horizon Index: KSP 63.33% ±0.00 90.50% ±0.00 27.17 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Fable 5.1 average GPT-6 Astra average
Legal 30.98% 22.42%
Finance 71.06% 62.86%
Healthcare 72.40% 68.20%
Math 100.00% 99.00%
Science 43.87% 58.00%
Education 48.53% 46.37%
Coding 54.79% 58.35%
Cyber 46.66% 48.97%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5.1 cost GPT-6 Astra cost Claude Fable 5.1 latency GPT-6 Astra latency
Vals Index $28.71 $18.46 1h17m 27m48s
Vals RSI Index $298.22 $759.51 90h00m 90h00m
Harvey's Legal Agent Benchmark $46.21 $26.16 1h52m 24m49s
Legal Research Bench $23.06 $10.58 57m29s 24m51s
EMB $15.93 $5.83 29m32s 13m24s
Finance Agent (v2) $8.35 $6.82 18m14s 11m33s
Tax Agent Bench $13.18 $5.80 28m34s 15m01s
MedCode N/A N/A 3m35s 102.86s
MedScribe N/A N/A 3m08s 2m34s
ProofBench v1.1 $2.70 $1.71 7m58s 5m08s
MysteryMechanism $5.63 $1.56 17m41s 8m09s
Terminal-Bench Science $38.01 $20.80 2h36m 1h49m
SAGE N/A N/A 105.75s 41.63s
Code Migration $70.97 $44.36 4h20m 58m30s
IOI $11.20 $6.50 30m06s 31m40s
ProgramBench $58.52 $11.57 2h27m 27m02s
Terminal-Bench 4.0 $17.18 $9.58 1h00m 35m42s
Vibe Code Bench 1-100 $149.66 $84.72 2h53m 3h26m
Vibe Code Bench v1.1 $33.37 $38.51 57m40s 39m04s
CyberBench v1.1 $4.01 $1.84 18m32s 4m10s
SRE Bench $32.23 $10.24 2h03m 17m51s
CUA-bench $323.78 $1821.37 37.44s 32.09s
Time Horizon Index: KSP $4655.01 $4206.59 N/A N/A

Results available only for Claude Fable 5.1

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SkillsBench
  • Public Benefits Bench v1.1

Results available only for GPT-6 Astra

  • BioMysteryBench
Model details Claude Fable 5.1 Model details GPT-6 Astra