Anthropic's Claude Fable 5 Tops Arena's Frontend Coding Benchmark With 72% Win Rate
Arena ranked Anthropic's Claude Fable 5 No. 1 in Code Arena: Frontend, saying the model won 72% of its frontend battles and led Opus-4.8 by 98 points. The model also ranked first in the HTML and React subleaderboards and in every listed frontend subcategory.
The frontend showing accompanies Claude Fable 5's No. 1 position on Agent Arena overall, where Arena earlier gave it a +11.2% net improvement score. Arena said that broader leaderboard was built from more than 300,000 tasks, 2 million tool calls and 40 million lines of code, with Claude Fable 5 leading in confirmed task success, praise versus complaint and tool hallucination, while placing 17th in steerability.
From the sources (14 posts)
@arenaICYMI: Agentic AI is now measured in the Arena. Agent Mode can handle deep research around competitive intelligence, market sizing & opportunity analysis, scientific & medical research and more. Every session shapes the Agent Arena leaderb
@ml_angelopoulosRT @arena: ICYMI: Agentic AI is now measured in the Arena. Agent Mode can handle deep research around competitive intelligence, market sizi…
@ml_angelopoulosIn case you didn’t notice: Agent Arena doesn’t have a voting mechanism. So how do we calculate the scores? The answer is causal inference. Agents are multi-stage systems where the orchestrator and harness work together to produce the end r
@petergostevRT @ml_angelopoulos: In case you didn’t notice: Agent Arena doesn’t have a voting mechanism. So how do we calculate the scores? The answer…
@arenaRT @ml_angelopoulos: In case you didn’t notice: Agent Arena doesn’t have a voting mechanism. So how do we calculate the scores? The answer…
@arenaGrok Build 0.1 ranks #15 and Grok 4.3 (High) #17 in the new Agent Arena leaderboard. Grok Build 0.1 improves meaningfully on bash capability over Grok 4.3. It is slightly less steerable and more prone to tool hallucinations, but looks to be
@arenaGrok Build 0.1 ranks #15 overall (-5.3%) - #15 Confirmed Success (-6.3%) - #18 Praise vs. Complaint (-15.8%) - #15 Steerability (-7.0%) - #9 Bash Recovery (+6.1%) - #19 Tool Hallucination (-3.5%)
@arenaGrok 4.3 (High) ranks #17 overall (-9.4%) - #20 Confirmed Success (-15.8%) - #19 Praise vs. Complaint (-16.6%) - #18 Steerability (-9.3%) - #16 Bash Recovery (-3.8%) - #17 Tool Hallucination (-1.6%)
@arenaExciting news: Claude Fable 5 ranks #1 on the new Agent Arena leaderboard! Fable 5 leads by the widest margin ever over Opus-4.8 and GPT-5.5 on two key signals: confirmed task success rate and praise vs. complaint, despite weaker steerabil
@arenaClaude Fable 5 by @AnthropicAI leads by the widest margins over other top models like Opus-4.8 and GPT-5.5 on two key signals: confirmed task success rate and praise vs. complaint.
@arenaClaude Fable 5 ranks #1 overall (+11.2%) - #1 Confirmed Task Success (+18.2%) - #1 Praise vs. Complaint (+30.6%) - #1 Tool Hallucination (+2.1%) - #7 Bash Recovery (+11.9%) - #17 Steerability (-6.8%, still stabilizing)
@scaling01insane jump in confirmed successes and praises by users
@arenaClaude Fable 5 ranks #1 in Code Arena: Frontend, leading by a wide margin over Opus-4.8. Highlights: - #1 in every sub leaderboard: HTML, React - #1 in every sub category: Brand & Marketing, Reference-Based Design, Data & Analytics, Consum
@arenaClaude Fable 5 ranks #1 in Code Arena: Frontend, winning 72% of its frontend battles and leading the arena by a wide +98 points.