Command Palette
Search for a command to run...

Anthropic's Claude Fable 5 Tops Epoch Capabilities Index at 161, 1 Point Ahead of GPT-5.5 Pro

aiai-modelingai-research-evals 6 posts · 4 accounts

Anthropic's Claude Fable 5 scored 161 on the Epoch Capabilities Index, moving to the top of the benchmark and edging GPT-5.5 Pro by one point. The result marks Anthropic's first lead on the ECI in more than a year.

Epoch said Fable 5 broke from Anthropic's earlier pattern of trailing the frontier overall while posting outsized gains mainly on software tasks, instead taking the lead on math benchmarks. It cited FrontierMath v2 results of 87% on Tiers 1–3 and 88% on Tier 4, while adding that there still is not enough data to determine whether Fable 5 also leads on software; WeirdML v2 alone would imply a software-focused ECI of 169.

From the sources (6 posts)

@epochairesearch

Claude Fable 5 achieves a new high score of 161 on the Epoch Capabilities Index! This beats out GPT-5.5 Pro by 1 point, and is the first time Anthropic has taken the lead on the ECI in over a year.

@epochairesearch

Historically Anthropic models have been slightly behind the frontier on the ECI, with outsized performance on software benchmarks compared to many other tasks. Fable 5 bucked this trend, taking the lead on math benchmarks:

@epochairesearch

There isn't yet enough data to say confidently whether Fable 5 also outperforms on software. It's performance on WeirdML v2 alone would suggest a SWE-specific ECI of 169.

@mtslive

SITUATION DETECTED: Claude Fable 5 has set a new record on the Epoch Capabilities Index, dethroning GPT-5.5 Pro. This is the first time an Anthropic model has been #1 on the benchmark since Claude 3.5 Sonnet in June 2024.

@andrewcurran_

RT @EpochAIResearch: Claude Fable 5 achieves a new high score of 161 on the Epoch Capabilities Index! This beats out GPT-5.5 Pro by 1 point…

@scaling01

Fable is slightly higher rated than GPT-5.5-Pro on Epoch's ECI I suspect as we get more benchmark results its ECI should improve to ~163

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive