Sep 1, 2026
Anthropic's Claude Fable 5.1 evaluated across our benchmark suite
We evaluated Anthropicโs new Claude Fable 5.1 across our benchmark suite.
-
Fable 5.1 takes #1 on the Vals Index (67.87%), ahead of Claude Opus 5 (67.21%) and its own predecessor Claude Fable 5 (66.04%).
-
It scores a perfect 100.00% on ProofBench v1.1, up from 95.00% for Fable 5 and matched only by the specialized prover AlephProver.
-
Coding results are the strongest part of the release: #1 on LiveCodeBench (90.52%), #2 on Terminal-Bench 2.1 (85.02%, up from 80.52% for Fable 5 and behind only GPT-5.6 Sol at 85.77%), and 90.26% on Vibe Code Bench, within noise of Fable 5โs leading 90.35%.
-
It also leads EMB (76.67%, up from 73.67%), MMLU Pro (92.38%), MMMU Pro (90.64%), and MedScribe (91.29%). On Legal Research Bench it matches Opus 5โs leading 55.29%, a 5.77-point gain over Fable 5.
-
Agentic legal work remains the weak spot: 6.67% on Harveyโs Legal Agent Benchmark (#18 of 55, with an 86.57% criteria pass rate), below Fable 5โs 11.25%. It also lands below the previous release on MedCode (53.51%, #7), SAGE (48.53%, #25), and TaxEval v2 (75.96%, #9).
Refusals remain frequent on bio- and cyber-adjacent tasks, so we ran Fable 5.1 with Claude Opus 5 and Claude Opus 4.8 as server-side fallbacks. Counting fallback-assisted tasks as failures changes Terminal-Bench 2.1 from 85.02% to 79.03% (23 of 267 tasks), the Vals Index from 67.87% to 66.85%, Legal Research Bench from 55.29% to 54.33%, and Harveyโs Legal Agent Benchmark from 6.67% to 5.83%. The effect is largest on ReverseEngBench, where 195 of 262 tasks (74.43%) were fallback-assisted and the score falls from 22.90% to 10.69%.
The model has a 1M-token context window and 128k max output tokens. Evaluations were run with compute effort set to โmaxโ and temperature set to 1.0.
Congrats to the Anthropic team on the release!