Anthropic’s Claude Opus 5 Tops Frontend Code and Text Arenas Under New Factuality Rank
Anthropic’s Claude Opus 5 model secured the number-one ranking on the Frontend Code and Text Arenas leaderboards after the Arena platform activated a factuality toggle merging human preference with verified accuracy. Using Max reasoning settings, the model led both competitive categories.
Arena noted that performance scores can fluctuate as they stabilize, and said a high-reasoning variant of the model also placed in the top three of both arenas. The rankings, based on agentic web coding and document reasoning tasks, marked the leading scores as preliminary pending convergence.
From the sources (25 posts)
@miroyatoRT @0xIlyy: Opus 5 cybersecurity requests fallback to Opus 4.8 too 😂
@artificialanlysClaude Opus 5 is the new leader on our agentic knowledge work benchmark, AA-Briefcase, outperforming Claude Fable 5 by nearly 150 Elo while reducing Cost per Task by 20% @AnthropicAI has released Claude Opus 5, the new leader on the Artifi
@perplexity_aiClaude Opus 5 is now available in Perplexity and Perplexity Computer. We evaluated it against six other models on WANDR. It outperformed all but Fable 5, while being 57% cheaper.
@reutersAnthropic rolls out Opus 5 AI model in efficiency upgrade
@andrew_n_carrRT @edis0n_zhang: We found that Opus 5 actually scores lower on FrontierCode 1.1 in xhigh and max reasoning efforts. Our investigation rev…
@artificialanlysClaude Opus 5 offers a wide range of intelligence-cost tradeoffs across its five effort settings. Max effort costs $17.79 per task, around 20% less than Claude Fable 5 ($22.30) while achieving a higher AA-Briefcase Elo. Notably, Claude Opus
@artificialanlysAA-Briefcase measures performance across objective rubric criteria, analytical quality, and presentation quality. Claude Opus 5’s gains are driven primarily by rubric pass rate and analytical quality. At max effort it achieves an Analytical
@artificialanlysClaude Opus 5’s top three effort settings all average more than 25 minutes per AA-Briefcase task (36.2, 34.3, and 25.7 minutes for max, xhigh, and high). At max effort this is roughly 50% longer than Opus 4.8 (24.1 minutes), primarily due t
@aravsrinivasRT @perplexity_ai: Claude Opus 5 is now available in Perplexity and Perplexity Computer. We evaluated it against six other models on WANDR…
@denisyaratsRT @perplexity_ai: Claude Opus 5 is now available in Perplexity and Perplexity Computer. We evaluated it against six other models on WANDR…
@cognitionRT @cl571128: We've received several questions about the Opus 5 FrontierCode results, where scores decline as reasoning effort increases. I…
@imjaredzRT @edis0n_zhang: We found that Opus 5 actually scores lower on FrontierCode 1.1 in xhigh and max reasoning efforts. Our investigation rev…
@polymarketJUST IN: Anthropic reveals Claude Opus 5 achieved a perfect 42/42 score on the 2026 International Mathematical Olympiad.
@teortaxestexOpenAI is still just much more efficient in hardcore reasoning
@scaling01Opus 5 ECI should be around 163.5 based on benchmarks they have provided their internal AECI is at 162.1, higher than the one of Mythos 5
@reutersAnthropic rolls out Opus 5 AI model in efficiency upgrade
@epochairesearchClaude Opus 5 achieves an ECI of 159, slightly below Fable 5's value of 161. However, looking only at software engineering benchmarks, we find that Opus 5 matches Fable 5's SWE-ECI of 161.
@scaling01Opus 5 ECI right now is at 159 ??? that seems incredibly underrated like literally 1 point better than Opus 4.8 while being much better at everything
@teortaxestexhow to shake Lisan's faith in any benchmark: show Anthropic doing meh on it (yes, ECI is all about 1 point diffs)
@scaling01@teortaxesTex well the score itself is fine. it's better than Opus 4.8 which makes sense but weaker than Fable it's more that the Anthropic ECI is so high that's confusing
@mtsliveSITUATION ANALYSIS: Claude Opus 5 After two months since Opus 4.8 and a month and a half since Fable 5, Claude Opus 5 has finally been released, and it seems to be a really good model, the best we’ve seen yet. For starters, it outperforms
@yacinemtbRT @herbiebradley: I dug into the Opus 5 system card & it does noticeably better on ARC-AGI 1 and 2 over 4.7/4.8. The description of solvi…
@juliusaiClaude Opus 5 is live in Julius. It joins both Fable 5 and Sonnet 5, and excels in both coding tasks and knowledge work in the Julius harness. From training and testing machine learning models; to conducting research across the web & build
@rohanpaul_aiRT @rohanpaul_ai: Anthropic launched Claude Opus 5 at half the price of its most powerful model Fable 5, while at the same time beating Feb…
@onusozSlopus is no more Claude Opus 4.8 -> 5 does a 20 point jump in AlmanBench and gets in the top 5 In the same league as GPT-5.6 Also 11 points higher than Sonnet 5 You know the performance gains are real (at least in multilingual singe-st