Moonshot Kimi K3 Leads Open Weights Field After Scoring 156 on Epoch Capabilities Index
Moonshot Kimi K3 achieved a new open weights record on the Epoch Capabilities Index with a score of 156, placing it just ahead of GPT 5.6 Luna and between Anthropic Opus 4.6 and OpenAI GPT 5.4.
The model also secured the top spot on the 3D Design Arena leaderboard with an Elo rating of 1,450, marking a 108 position improvement. Predictions for the upcoming release of GPT-6 place its index score at 166.5, implying U.S. frontier laboratories retain a lead of roughly eight to 11 months.
From the sources (16 posts)
@designarenaBREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this open-weight model leads the way, ahead of Fable 5, Sonnet 5, and Opus 4.8 by @AnthropicAI Huge congrats to the @Kimi
@kimi_moonshotRT @DesignArena: BREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this…
@crystalsssupKimi is #1 in Design Arena, a benchmark for frontend website-building capabilities.
@arenaExciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a maj
@arenaKimi K3 ranks #4 overall (+9.6%) - #1 Confirmed Task Success (+14.4%) - #3 Praise vs. Complaint (+20.6%) - #4 Tool Hallucination (+1.1%) - #14 Steerability (+5.6%) - #17 Bash Recovery (+6.4%)
@scaling01RT @arena: Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's…
@kimi_moonshotRT @arena: Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's…
@rohanpaul_aiKimi K3 is ahead of Claude Feble 5 again. Has taken the top spot on DesignArena's Frontend Web App benchmark. On DesignArena bench, AI models receive the same app-building prompt, produce competing interfaces, and users vote for the bette
@epochairesearchMoonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it between Opus 4.6, and GPT 5.4, which released in February and March 2026 respectively, and just ahead of GPT 5.6 Luna. ht
@scaling01RT @EpochAIResearch: Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it…
@teortaxestexKimi K3 scores 5'11 3/4" on Epoch Capabilities index, ahead of GPT 5.6 Luna at 5.11" GPT 5.6 Terra (Gigachad) reigns supreme at 6' (actually 155.6 vs 155.3 vs 158.4) @scaling01 takes the W
@zacharynadoRT @peterwildeford: Kimi K3 is exactly the level of capability you would predict it to have given the long-term 2 year trend of Chinese AI…
@kimi_moonshotRT @DesignArena: BREAKING: Kimi K3 by @Kimi_Moonshot is 1st overall on 3D Design with an Elo of 1450. This is a 6 position and 108 Elo jum…
@crystalsssupKimi K3 is #1 in 3D design arena
@teortaxestexThis is absurdly impressive. BrokenArXiv is *hard* for LLMs and OpenAI focused on such problems hard and holds a commanding lead. Kimi is not a frontend slop machine, it's a generalist proto-AGI (though same can be said of Meta, congrats) h
@scaling01my prediction for GPT-6's ECI is 166.5 (given that it's released at the end of august) so roughly ~10 ECI points higher than Kimi-K3 depending on which trendline you take this implies US frontier labs are ahead 8.2 - 10.9 months