Claude Fable 5 vs Gemini 3.8 Flash: Benchmark Comparison
Claude Fable 5 has the higher score on 23 of 25 shared benchmarks; Gemini 3.8 Flash leads on 2.
The largest observed score gap is 47.00 pts on ProofBench v1.1 , where Claude Fable 5 leads.
Reported ±1 standard-error ranges overlap on 8 of 24 shared benchmarks with comparable uncertainty data. This is not a pairwise statistical significance test.
Shared benchmark results
Scores and reported standard errors come from the latest visible Vals benchmark version available for both models.
| Benchmark | Claude Fable 5 | Gemini 3.8 Flash | Gap | Reported uncertainty |
|---|---|---|---|---|
| Vals Index | 61.39% ±1.00 | 54.83% ±1.03 | 6.56 pts | Reported ±1 SE ranges do not overlap |
| Vals RSI Index | 24.32% | 19.93% | 4.39 pts | Uncertainty comparison unavailable |
| Harvey's Legal Agent Benchmark | 11.25% ±2.15 | 10.00% ±2.42 | 1.25 pts | Reported ±1 SE ranges overlap |
| Legal Research Bench | 49.52% ±3.48 | 38.94% ±3.39 | 10.58 pts | Reported ±1 SE ranges do not overlap |
| LegalBench | 88.56% ±0.33 | 86.99% ±0.43 | 1.57 pts | Reported ±1 SE ranges do not overlap |
| EMB | 73.67% ±2.47 | 72.20% ±2.42 | 1.47 pts | Reported ±1 SE ranges overlap |
| Finance Agent (v2) | 56.31% ±0.84 | 61.44% ±0.13 | 5.12 pts | Reported ±1 SE ranges do not overlap |
| MortgageTax | 68.92% ±0.91 | 65.34% ±0.94 | 3.58 pts | Reported ±1 SE ranges do not overlap |
| Tax Agent Bench | 69.82% ±3.12 | 66.77% ±3.15 | 3.05 pts | Reported ±1 SE ranges overlap |
| TaxEval v2 | 76.94% ±0.82 | 74.45% ±0.85 | 2.49 pts | Reported ±1 SE ranges do not overlap |
| MedCode | 56.07% ±2.20 | 48.13% ±2.18 | 7.94 pts | Reported ±1 SE ranges do not overlap |
| MedScribe | 88.52% ±1.95 | 84.50% ±1.94 | 4.03 pts | Reported ±1 SE ranges do not overlap |
| ProofBench v1.1 | 95.00% ±2.19 | 48.00% ±5.02 | 47.00 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench Science | 15.71% ±4.38 | 8.57% ±3.37 | 7.14 pts | Reported ±1 SE ranges overlap |
| GPQA Diamond | 93.18% ±1.94 | 94.44% ±1.48 | 1.26 pts | Reported ±1 SE ranges overlap |
| MMLU Pro | 91.50% ±0.28 | 90.22% ±0.29 | 1.28 pts | Reported ±1 SE ranges do not overlap |
| MMMU Pro | 89.31% ±0.74 | 89.08% ±0.75 | 0.23 pts | Reported ±1 SE ranges overlap |
| SAGE | 51.89% ±3.40 | 35.06% ±3.36 | 16.83 pts | Reported ±1 SE ranges do not overlap |
| Code Migration | 55.06% ±4.61 | 36.55% ±4.18 | 18.52 pts | Reported ±1 SE ranges do not overlap |
| LiveCodeBench | 89.78% ±0.89 | 89.48% ±0.90 | 0.29 pts | Reported ±1 SE ranges overlap |
| ProgramBench | 2.00% ±0.99 | 1.00% ±0.70 | 1.00 pts | Reported ±1 SE ranges overlap |
| SWE-bench | 95.00% ±0.98 | 80.00% ±1.79 | 15.00 pts | Reported ±1 SE ranges do not overlap |
| Terminal-Bench 4.0 | 41.41% ±1.34 | 19.19% ±2.52 | 22.22 pts | Reported ±1 SE ranges do not overlap |
| Vibe Code Bench v1.1 | 90.35% ±2.10 | 78.65% ±3.88 | 11.70 pts | Reported ±1 SE ranges do not overlap |
| Public Benefits Bench v1.1 | 70.43% ±1.19 | 65.29% ±1.24 | 5.14 pts | Reported ±1 SE ranges do not overlap |
Performance by category
| Category | Claude Fable 5 average | Gemini 3.8 Flash average |
|---|---|---|
| Legal | 49.78% | 45.31% |
| Finance | 69.13% | 68.04% |
| Healthcare | 72.30% | 66.32% |
| Math | 95.00% | 48.00% |
| Science | 15.71% | 8.57% |
| Academic | 91.33% | 91.25% |
| Education | 51.89% | 35.06% |
| Coding | 62.27% | 50.81% |
| Social Mobility | 70.43% | 65.29% |
Cost and latency
Cost per test appears only where the benchmark reports it for both models. Latency is the measured completion time for that benchmark.
| Benchmark | Claude Fable 5 cost | Gemini 3.8 Flash cost | Claude Fable 5 latency | Gemini 3.8 Flash latency |
|---|---|---|---|---|
| Vals Index | $29.59 | $5.73 | 42m24s | 51m57s |
| Vals RSI Index | $491.22 | $332.04 | 90h00m | 90h00m |
| Harvey's Legal Agent Benchmark | $19.23 | $3.66 | 26m53s | 29m21s |
| Legal Research Bench | $9.79 | $1.63 | 22m18s | 5m24s |
| LegalBench | N/A | N/A | 8.96s | 3.32s |
| EMB | $12.35 | $8.24 | 26m34s | 12m49s |
| Finance Agent (v2) | $8.06 | $2.00 | 10m11s | 3m21s |
| MortgageTax | N/A | N/A | 16.10s | 14.79s |
| Tax Agent Bench | $6.70 | $0.86 | 14m27s | 2m58s |
| TaxEval v2 | N/A | N/A | 56.87s | 8.76s |
| MedCode | N/A | N/A | 91.44s | 42.89s |
| MedScribe | N/A | N/A | 119.47s | 21.14s |
| ProofBench v1.1 | $5.45 | $0.60 | 13m43s | 6m09s |
| Terminal-Bench Science | $51.99 | $5.64 | 1h59m | 55m11s |
| GPQA Diamond | N/A | N/A | 99.90s | 20.18s |
| MMLU Pro | N/A | N/A | 25.00s | 7.82s |
| MMMU Pro | N/A | N/A | 61.44s | 12.03s |
| SAGE | N/A | N/A | 116.92s | 28.64s |
| Code Migration | $112.10 | $18.49 | 1h49m | 2h37m |
| LiveCodeBench | N/A | N/A | 118.53s | 24.07s |
| ProgramBench | $75.68 | $10.93 | 2h37m | 46m39s |
| SWE-bench | $2.05 | $2.19 | 5m56s | 11m23s |
| Terminal-Bench 4.0 | $30.34 | $8.77 | 1h08m | 1h48m |
| Vibe Code Bench v1.1 | $41.71 | $6.87 | 1h02m | 8m39s |
| Public Benefits Bench v1.1 | $4.59 | $1.01 | 21m17s | 3m21s |
Results available only for Claude Fable 5
None.
Results available only for Gemini 3.8 Flash
- BioMysteryBench
- MysteryMechanism
- IOI
- SkillsBench
- Vibe Code Bench 1-100
- CyberBench v1.1
- CUA-bench
- Time Horizon Index: KSP