Command Palette
Search for a command to run...

Anthropic's Claude Fable 5 Takes Agent Arena Lead at +11.2%, Overtakes GPT-5.5

aiai-modelingai-research-evalsai-productsai-agents-coding 12 posts · 4 accounts

Anthropic's Claude Fable 5 has taken the top spot on Agent Arena, a benchmark for long-horizon AI agent tasks, with an overall net improvement score of +11.2%, according to Arena. The model ranked first in confirmed task success (+18.2%), praise versus complaint (+30.6%) and tool hallucination (+2.1%), and Arena said it opened the widest lead yet over Anthropic's Opus-4.8 and OpenAI's GPT-5.5 on the task-success and praise measures.

The update displaces OpenAI's GPT-5.5, which led Agent Arena when the leaderboard debuted on June 8. Arena says the system measures millions of live sessions in which models use web search, filesystem and terminal tools, with the snapshot built from more than 300,000 tasks, 2 million tool calls and 40 million lines of code and scored through a causal-tracing method across five signals. Claude Fable 5's weaker area was steerability, where it ranked 17th at -6.8%, while placing seventh in bash recovery at +11.9%.

From the sources (12 posts)

@arena

ICYMI: Agentic AI is now measured in the Arena. Agent Mode can handle deep research around competitive intelligence, market sizing & opportunity analysis, scientific & medical research and more. Every session shapes the Agent Arena leaderb

@ml_angelopoulos

RT @arena: ICYMI: Agentic AI is now measured in the Arena. Agent Mode can handle deep research around competitive intelligence, market sizi…

@ml_angelopoulos

In case you didn’t notice: Agent Arena doesn’t have a voting mechanism. So how do we calculate the scores? The answer is causal inference. Agents are multi-stage systems where the orchestrator and harness work together to produce the end r

@petergostev

RT @ml_angelopoulos: In case you didn’t notice: Agent Arena doesn’t have a voting mechanism. So how do we calculate the scores? The answer…

@arena

RT @ml_angelopoulos: In case you didn’t notice: Agent Arena doesn’t have a voting mechanism. So how do we calculate the scores? The answer…

@arena

Grok Build 0.1 ranks #15 and Grok 4.3 (High) #17 in the new Agent Arena leaderboard. Grok Build 0.1 improves meaningfully on bash capability over Grok 4.3. It is slightly less steerable and more prone to tool hallucinations, but looks to be

@arena

Grok Build 0.1 ranks #15 overall (-5.3%) - #15 Confirmed Success (-6.3%) - #18 Praise vs. Complaint (-15.8%) - #15 Steerability (-7.0%) - #9 Bash Recovery (+6.1%) - #19 Tool Hallucination (-3.5%)

@arena

Grok 4.3 (High) ranks #17 overall (-9.4%) - #20 Confirmed Success (-15.8%) - #19 Praise vs. Complaint (-16.6%) - #18 Steerability (-9.3%) - #16 Bash Recovery (-3.8%) - #17 Tool Hallucination (-1.6%)

@arena

Exciting news: Claude Fable 5 ranks #1 on the new Agent Arena leaderboard! Fable 5 leads by the widest margin ever over Opus-4.8 and GPT-5.5 on two key signals: confirmed task success rate and praise vs. complaint, despite weaker steerabil

@arena

Claude Fable 5 by @AnthropicAI leads by the widest margins over other top models like Opus-4.8 and GPT-5.5 on two key signals: confirmed task success rate and praise vs. complaint.

@arena

Claude Fable 5 ranks #1 overall (+11.2%) - #1 Confirmed Task Success (+18.2%) - #1 Praise vs. Complaint (+30.6%) - #1 Tool Hallucination (+2.1%) - #7 Bash Recovery (+11.9%) - #17 Steerability (-6.8%, still stabilizing)

@scaling01

insane jump in confirmed successes and praises by users

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive