Command Palette
Search for a command to run...

Moonshot’s Kimi K3 Model Tops Frontend Web Benchmark and Ranks Fourth on Agent Arena Leaderboard, Matching Claude and GPT

aiai-modelingai-research-evals 19 posts · 14 accounts

Moonshot AI’s Kimi K3 model placed fourth on the Agent Arena leaderboard, matching results from Claude Opus 4.8 and GPT-5.6 Sol. The ranking, based on more than 8,000 live sessions, shows the model leading in confirmed task success and posting a stronger praise-to-complaint ratio than most competitors.

The model also secured the top spot on DesignArena’s Frontend Web App benchmark with an Elo score of 1,326, outperforming Fable 5 and Sonnet 5. Weights for the 2.8 trillion parameter model are scheduled for release on July 27. Agent Arena evaluates AI systems using web search, filesystem, and terminal tools to complete long-horizon workflows, though Kimi K3 currently scores lower on steerability and bash recovery.

From the sources (19 posts)

@fs0c131y

RT @Kimi_Moonshot: Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi…

@zeyuanallenzhu

Congrats, @Kimi_Moonshot! 『 Kimi’s Four Commandments』have circulated in the Chinese AI community for months. Many people add their own fifth for comic relief, but the original four were the core. Here's an English translation --- since appa

@vipulved

Kimi K3 reduces frontier AI costs by 3x. Excited to serve this model natively on @togethercompute starting July 27th!

@crystalsssup

RT @AfterQuery: Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5. An open weight model now outperforms al…

@nandodf

RT @SemiAnalysis_: CHINA’S KIMI K3 HAS SURPASSED ALL AMERICAN MODELS IN FRONT-END CODING WHILE BEING SMALLER THAN MOST CLOSED-SOURCE FRONTI…

@crystalsssup

RT @levelsio: Kimi K3 is absolutely hammering through my Windows XP Simulator to do list Claude Code couldn't do this for 2 weeks or kept…

@cline

Kimi K3 is now on ClinePass! 🚀 $9.99/month subscription for 2-5x discounted access to it and other open weight models like GLM, DeepSeek, and others. Use it on Cline CLI & extension with $1.99 special promo if sign up via: npm i -g cl

@hrkrshnn

@DavidSacks Yes, you can even see Anthropic's distilled reasoning traces inside Kimi on cyber tasks.

@mikebelshe

Just wait til the regulators come help

@calilyliu

RT @DavidSacks: This is concerning. For the first time, a Chinese model Kimi K3 has taken #1 on the Frontend Code Arena and is scoring at o…

@rohanpaul_ai

Kimi K3 may be super useful for finance, planning and operations, given how well it plays with spreadsheets. Took the top spot on SpreadsheetBench 2, almost tallying with Fable 5 and completing 34.8% of workflows. SpreadsheetBench 2 tests

@designarena

BREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this open-weight model leads the way, ahead of Fable 5, Sonnet 5, and Opus 4.8 by @AnthropicAI Huge congrats to the @Kimi

@kimi_moonshot

RT @DesignArena: BREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this…

@crystalsssup

Kimi is #1 in Design Arena, a benchmark for frontend website-building capabilities.

@arena

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a maj

@arena

Kimi K3 ranks #4 overall (+9.6%) - #1 Confirmed Task Success (+14.4%) - #3 Praise vs. Complaint (+20.6%) - #4 Tool Hallucination (+1.1%) - #14 Steerability (+5.6%) - #17 Bash Recovery (+6.4%)

@scaling01

RT @arena: Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's…

@kimi_moonshot

RT @arena: Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's…

@rohanpaul_ai

Kimi K3 is ahead of Claude Feble 5 again. Has taken the top spot on DesignArena's Frontend Web App benchmark. On DesignArena bench, AI models receive the same app-building prompt, produce competing interfaces, and users vote for the bette

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive