Claude Fable 5 vs Gemini 3.8 Flash: Benchmark Comparison

Claude Fable 5 has the higher score on 23 of 25 shared benchmarks; Gemini 3.8 Flash leads on 2.

The largest observed score gap is 47.00 pts on ProofBench v1.1 , where Claude Fable 5 leads.

Reported ±1 standard-error ranges overlap on 8 of 24 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.

Shared benchmark results

Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.

Benchmark Claude Fable 5 Gemini 3.8 Flash Gap Reported uncertainty
Vals Index 61.39% ±1.00 54.83% ±1.03 6.56 pts Reported ±1 SE ranges do not overlap
Vals RSI Index 24.32% 19.93% 4.39 pts Uncertainty comparison unavailable
Harvey's Legal Agent Benchmark 11.25% ±2.15 10.00% ±2.42 1.25 pts Reported ±1 SE ranges overlap
Legal Research Bench 49.52% ±3.48 38.94% ±3.39 10.58 pts Reported ±1 SE ranges do not overlap
LegalBench 88.56% ±0.33 86.99% ±0.43 1.57 pts Reported ±1 SE ranges do not overlap
EMB 73.67% ±2.47 72.20% ±2.42 1.47 pts Reported ±1 SE ranges overlap
Finance Agent (v2) 56.31% ±0.84 61.44% ±0.13 5.12 pts Reported ±1 SE ranges do not overlap
MortgageTax 68.92% ±0.91 65.34% ±0.94 3.58 pts Reported ±1 SE ranges do not overlap
Tax Agent Bench 69.82% ±3.12 66.77% ±3.15 3.05 pts Reported ±1 SE ranges overlap
TaxEval v2 76.94% ±0.82 74.45% ±0.85 2.49 pts Reported ±1 SE ranges do not overlap
MedCode 56.07% ±2.20 48.13% ±2.18 7.94 pts Reported ±1 SE ranges do not overlap
MedScribe 88.52% ±1.95 84.50% ±1.94 4.03 pts Reported ±1 SE ranges do not overlap
ProofBench v1.1 95.00% ±2.19 48.00% ±5.02 47.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench Science 15.71% ±4.38 8.57% ±3.37 7.14 pts Reported ±1 SE ranges overlap
GPQA Diamond 93.18% ±1.94 94.44% ±1.48 1.26 pts Reported ±1 SE ranges overlap
MMLU Pro 91.50% ±0.28 90.22% ±0.29 1.28 pts Reported ±1 SE ranges do not overlap
MMMU Pro 89.31% ±0.74 89.08% ±0.75 0.23 pts Reported ±1 SE ranges overlap
SAGE 51.89% ±3.40 35.06% ±3.36 16.83 pts Reported ±1 SE ranges do not overlap
Code Migration 55.06% ±4.61 36.55% ±4.18 18.52 pts Reported ±1 SE ranges do not overlap
LiveCodeBench 89.78% ±0.89 89.48% ±0.90 0.29 pts Reported ±1 SE ranges overlap
ProgramBench 2.00% ±0.99 1.00% ±0.70 1.00 pts Reported ±1 SE ranges overlap
SWE-bench 95.00% ±0.98 80.00% ±1.79 15.00 pts Reported ±1 SE ranges do not overlap
Terminal-Bench 4.0 41.41% ±1.34 19.19% ±2.52 22.22 pts Reported ±1 SE ranges do not overlap
Vibe Code Bench v1.1 90.35% ±2.10 78.65% ±3.88 11.70 pts Reported ±1 SE ranges do not overlap
Public Benefits Bench v1.1 70.43% ±1.19 65.29% ±1.24 5.14 pts Reported ±1 SE ranges do not overlap

Performance by category

Category Claude Fable 5 average Gemini 3.8 Flash average
Legal 49.78% 45.31%
Finance 69.13% 68.04%
Healthcare 72.30% 66.32%
Math 95.00% 48.00%
Science 15.71% 8.57%
Academic 91.33% 91.25%
Education 51.89% 35.06%
Coding 62.27% 50.81%
Social Mobility 70.43% 65.29%

Cost and latency

Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.

Benchmark Claude Fable 5 cost Gemini 3.8 Flash cost Claude Fable 5 latency Gemini 3.8 Flash latency
Vals Index $29.59 $5.73 42m24s 51m57s
Vals RSI Index $491.22 $332.04 90h00m 90h00m
Harvey's Legal Agent Benchmark $19.23 $3.66 26m53s 29m21s
Legal Research Bench $9.79 $1.63 22m18s 5m24s
LegalBench N/A N/A 8.96s 3.32s
EMB $12.35 $8.24 26m34s 12m49s
Finance Agent (v2) $8.06 $2.00 10m11s 3m21s
MortgageTax N/A N/A 16.10s 14.79s
Tax Agent Bench $6.70 $0.86 14m27s 2m58s
TaxEval v2 N/A N/A 56.87s 8.76s
MedCode N/A N/A 91.44s 42.89s
MedScribe N/A N/A 119.47s 21.14s
ProofBench v1.1 $5.45 $0.60 13m43s 6m09s
Terminal-Bench Science $51.99 $5.64 1h59m 55m11s
GPQA Diamond N/A N/A 99.90s 20.18s
MMLU Pro N/A N/A 25.00s 7.82s
MMMU Pro N/A N/A 61.44s 12.03s
SAGE N/A N/A 116.92s 28.64s
Code Migration $112.10 $18.49 1h49m 2h37m
LiveCodeBench N/A N/A 118.53s 24.07s
ProgramBench $75.68 $10.93 2h37m 46m39s
SWE-bench $2.05 $2.19 5m56s 11m23s
Terminal-Bench 4.0 $30.34 $8.77 1h08m 1h48m
Vibe Code Bench v1.1 $41.71 $6.87 1h02m 8m39s
Public Benefits Bench v1.1 $4.59 $1.01 21m17s 3m21s

Results available only for Claude Fable 5

None.

Results available only for Gemini 3.8 Flash

  • BioMysteryBench
  • MysteryMechanism
  • IOI
  • SkillsBench
  • Vibe Code Bench 1-100
  • CyberBench v1.1
  • CUA-bench
  • Time Horizon Index: KSP
Model details Claude Fable 5 Model details Gemini 3.8 Flash