Claude Fable 5 vs GPT-6 Astra: Benchmark Comparison

Claude Fable 5 has the higher score on 9 of 16 shared benchmarks; GPT-6 Astra leads on 7.

The largest observed score gap is 47.14 pts on Terminal-Bench Science , where GPT-6 Astra leads.

Reported ±1 standard-error ranges overlap on 6 of 15 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5 GPT-6 Astra Gap Reported uncertainty
Vals Index 61.39% ±1.00 63.13% ±1.19 1.74 pts Reported ±1 SE ranges overlap
Vals RSI Index 24.32% 27.07% 2.75 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 11.25% ±2.15 5.42% ±1.17 5.83 pts Reported ±1 SE ranges do not overlap
Legal Research Bench 49.52% ±3.48 39.42% ±3.40 10.10 pts Reported ±1 SE ranges do not overlap
EMB 73.67% ±2.47 71.70% ±2.48 1.97 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 56.31% ±0.84 53.54% ±2.08 2.77 pts Reported ±1 SE ranges overlap
Tax Agent Bench 69.82% ±3.12 63.34% ±3.15 6.49 pts Reported ±1 SE ranges do not overlap
MedCode 56.07% ±2.20 48.49% ±2.13 7.58 pts Reported ±1 SE ranges do not overlap
MedScribe 88.52% ±1.95 87.91% ±1.94 0.61 pts Reported ±1 SE ranges overlap
ProofBench v1.1 95.00% ±2.19 99.00% ±1.00 4.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 15.71% ±4.38 62.86% ±5.82 47.14 pts Reported ±1 SE ranges do not overlap
SAGE 51.89% ±3.40 46.37% ±3.44 5.52 pts Reported ±1 SE ranges overlap
Code Migration 55.06% ±4.61 67.74% ±4.22 12.67 pts Reported ±1 SE ranges do not overlap
ProgramBench 2.00% ±0.99 5.50% ±1.62 3.50 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 41.41% ±1.34 59.60% ±4.40 18.18 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 90.35% ±2.10 89.59% ±2.17 0.76 pts Reported ±1 SE ranges overlap

Performance by category

Category Claude Fable 5 average GPT-6 Astra average
Legal 30.38% 22.42%
Finance 66.60% 62.86%
Healthcare 72.30% 68.20%
Math 95.00% 99.00%
Science 15.71% 62.86%
Education 51.89% 46.37%
Coding 47.21% 55.61%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5 cost GPT-6 Astra cost Claude Fable 5 latency GPT-6 Astra latency
Vals Index $29.59 $18.46 42m24s 27m48s
Vals RSI Index $491.22 $759.51 90h00m 90h00m
Harvey's Legal Agent Benchmark $19.23 $26.16 26m53s 24m49s
Legal Research Bench $9.79 $10.58 22m18s 24m51s
EMB $12.35 $5.83 26m34s 13m24s
Finance Agent (v2) $8.06 $6.82 10m11s 11m33s
Tax Agent Bench $6.70 $5.80 14m27s 15m01s
MedCode N/A N/A 91.44s 102.86s
MedScribe N/A N/A 119.47s 2m34s
ProofBench v1.1 $5.45 $1.71 13m43s 5m08s
Terminal-Bench Science $51.99 $20.80 1h59m 1h49m
SAGE N/A N/A 116.92s 41.63s
Code Migration $112.10 $44.36 1h49m 58m30s
ProgramBench $75.68 $11.57 2h37m 27m02s
Terminal-Bench 4.0 $30.34 $9.58 1h08m 35m42s
Vibe Code Bench v1.1 $41.71 $38.51 1h02m 39m04s

Results available only for Claude Fable 5

  • LegalBench
  • MortgageTax
  • TaxEval v2
  • GPQA Diamond
  • MMLU Pro
  • MMMU Pro
  • LiveCodeBench
  • SWE-bench
  • Public Benefits Bench v1.1

Results available only for GPT-6 Astra

  • BioMysteryBench
  • MysteryMechanism
  • IOI
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • SRE Bench
  • CUA-bench
  • Time Horizon Index: KSP
Model details Claude Fable 5 Model details GPT-6 Astra