Moonshot’s Kimi K3 Model Tops Frontend Web Benchmark and Ranks Fourth on Agent Arena Leaderboard, Matching Claude and GPT
Moonshot AI’s Kimi K3 model placed fourth on the Agent Arena leaderboard, matching results from Claude Opus 4.8 and GPT-5.6 Sol. The ranking, based on more than 8,000 live sessions, shows the model leading in confirmed task success and posting a stronger praise-to-complaint ratio than most competitors.
The model also secured the top spot on DesignArena’s Frontend Web App benchmark with an Elo score of 1,326, outperforming Fable 5 and Sonnet 5. Weights for the 2.8 trillion parameter model are scheduled for release on July 27. Agent Arena evaluates AI systems using web search, filesystem, and terminal tools to complete long-horizon workflows, though Kimi K3 currently scores lower on steerability and bash recovery.
From the sources (19 posts)
@fs0c131yRT @Kimi_Moonshot: Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi…
@zeyuanallenzhuCongrats, @Kimi_Moonshot! 『 Kimi’s Four Commandments』have circulated in the Chinese AI community for months. Many people add their own fifth for comic relief, but the original four were the core. Here's an English translation --- since appa
@vipulvedKimi K3 reduces frontier AI costs by 3x. Excited to serve this model natively on @togethercompute starting July 27th!
@crystalsssupRT @AfterQuery: Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5. An open weight model now outperforms al…
@nandodfRT @SemiAnalysis_: CHINA’S KIMI K3 HAS SURPASSED ALL AMERICAN MODELS IN FRONT-END CODING WHILE BEING SMALLER THAN MOST CLOSED-SOURCE FRONTI…
@crystalsssupRT @levelsio: Kimi K3 is absolutely hammering through my Windows XP Simulator to do list Claude Code couldn't do this for 2 weeks or kept…
@clineKimi K3 is now on ClinePass! 🚀 $9.99/month subscription for 2-5x discounted access to it and other open weight models like GLM, DeepSeek, and others. Use it on Cline CLI & extension with $1.99 special promo if sign up via: npm i -g cl
@hrkrshnn@DavidSacks Yes, you can even see Anthropic's distilled reasoning traces inside Kimi on cyber tasks.
@mikebelsheJust wait til the regulators come help
@calilyliuRT @DavidSacks: This is concerning. For the first time, a Chinese model Kimi K3 has taken #1 on the Frontend Code Arena and is scoring at o…
@rohanpaul_aiKimi K3 may be super useful for finance, planning and operations, given how well it plays with spreadsheets. Took the top spot on SpreadsheetBench 2, almost tallying with Fable 5 and completing 34.8% of workflows. SpreadsheetBench 2 tests
@designarenaBREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this open-weight model leads the way, ahead of Fable 5, Sonnet 5, and Opus 4.8 by @AnthropicAI Huge congrats to the @Kimi
@kimi_moonshotRT @DesignArena: BREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this…
@crystalsssupKimi is #1 in Design Arena, a benchmark for frontend website-building capabilities.
@arenaExciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a maj
@arenaKimi K3 ranks #4 overall (+9.6%) - #1 Confirmed Task Success (+14.4%) - #3 Praise vs. Complaint (+20.6%) - #4 Tool Hallucination (+1.1%) - #14 Steerability (+5.6%) - #17 Bash Recovery (+6.4%)
@scaling01RT @arena: Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's…
@kimi_moonshotRT @arena: Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's…
@rohanpaul_aiKimi K3 is ahead of Claude Feble 5 again. Has taken the top spot on DesignArena's Frontend Web App benchmark. On DesignArena bench, AI models receive the same app-building prompt, produce competing interfaces, and users vote for the bette