Command Palette
Search for a command to run...

Anthropic’s Claude Opus 5 Max Model Ranks No. 2 on Agent Arena With 11.88% Net Improvement

aiai-modelingai-research-evals 5 posts · 2 accounts

Anthropic’s Claude Opus 5 Max configuration ranks second on the Agent Arena benchmark following more than 7,000 real-world agentic sessions, posting a net improvement of 11.88%.

The default Opus 5 High model secured third place with an 11.73% improvement, placing it behind Fable 5 and ahead of GPT-5.6 Sol. The benchmark evaluates long-horizon tasks using a causal tracing methodology, and the operator noted the Max configuration’s score remains preliminary and subject to further convergence tracking.

From the sources (5 posts)

@scaling01

RT @arena: Exciting news: Claude Opus 5 with Max reasoning is #1 in the Frontend Code Arena and Text Arena with factuality on! Claude Opus…

@arena

In the Text Arena, Claude Opus 5 with Max reasoning ranks #1 with factuality on. Factuality is a new ranking that combines human preference with factual accuracy. We audit battles by sampling responses, extracting verifiable claims, and c

@arena

Exciting news: Claude Opus 5 with Max reasoning is #1 in the Frontend Code Arena and Text Arena with factuality on! Claude Opus 5 with default reasoning high is also very strong landing #3 in Frontend Code Arena, right behind Kimi K3 - and

@arena

Exciting news: @AnthropicAI's Claude Opus 5 (Max) is #2 in Agent Arena, with Opus 5 (High) right behind at #3, based on over 7K real-world agentic sessions. A strong debut: it slots in just below #1 Fable 5, and ahead of GPT-5.6 Sol (xHigh)

@scaling01

RT @arena: Exciting news: @AnthropicAI's Claude Opus 5 (Max) is #2 in Agent Arena, with Opus 5 (High) right behind at #3, based on over 7K…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive