Command Palette
Search for a command to run...

Moonshot Kimi K3 Scores Second on Agentic Benchmark Behind Claude, Costs $10.57 Per Task

aiai-modelingai-research-evals 21 posts · 10 accounts

Moonshot AI’s Kimi K3 model scored an Elo rating of 1543 on Artificial Analysis’s new AA-Briefcase benchmark for agentic knowledge work, finishing second overall behind Anthropic’s Claude Fable 5, which scored 1574. The 2.8 trillion parameter model outperformed GPT-5.6 Sol and Claude Opus 4.8, recording a 727 point improvement over its predecessor, Kimi K2.6.

The benchmark tests models on complex tasks using an internal private dataset and combines correctness, analytical quality, and presentation quality into a single metric. Kimi K3 recorded a strong analytical quality score comparable to Fable 5 but a weaker presentation quality rating. The model averages $10.57 per task and requires 56.4 minutes to complete a single assignment, making it roughly twice as slow and significantly more expensive than leading rivals. Its model weights are scheduled for public release on July 27.

From the sources (21 posts)

@designarena

BREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this open-weight model leads the way, ahead of Fable 5, Sonnet 5, and Opus 4.8 by @AnthropicAI Huge congrats to the @Kimi

@kimi_moonshot

RT @DesignArena: BREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena With an Elo of 1326, this…

@crystalsssup

Kimi is #1 in Design Arena, a benchmark for frontend website-building capabilities.

@arena

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open-weight model. This release marks a maj

@arena

Kimi K3 ranks #4 overall (+9.6%) - #1 Confirmed Task Success (+14.4%) - #3 Praise vs. Complaint (+20.6%) - #4 Tool Hallucination (+1.1%) - #14 Steerability (+5.6%) - #17 Bash Recovery (+6.4%)

@scaling01

RT @arena: Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's…

@kimi_moonshot

RT @arena: Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's…

@rohanpaul_ai

Kimi K3 is ahead of Claude Feble 5 again. Has taken the top spot on DesignArena's Frontend Web App benchmark. On DesignArena bench, AI models receive the same app-building prompt, produce competing interfaces, and users vote for the bette

@epochairesearch

Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it between Opus 4.6, and GPT 5.4, which released in February and March 2026 respectively, and just ahead of GPT 5.6 Luna. ht

@scaling01

RT @EpochAIResearch: Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it…

@teortaxestex

Kimi K3 scores 5'11 3/4" on Epoch Capabilities index, ahead of GPT 5.6 Luna at 5.11" GPT 5.6 Terra (Gigachad) reigns supreme at 6' (actually 155.6 vs 155.3 vs 158.4) @scaling01 takes the W

@zacharynado

RT @peterwildeford: Kimi K3 is exactly the level of capability you would predict it to have given the long-term 2 year trend of Chinese AI…

@kimi_moonshot

RT @DesignArena: BREAKING: Kimi K3 by @Kimi_Moonshot is 1st overall on 3D Design with an Elo of 1450. This is a 6 position and 108 Elo jum…

@crystalsssup

Kimi K3 is #1 in 3D design arena

@teortaxestex

This is absurdly impressive. BrokenArXiv is *hard* for LLMs and OpenAI focused on such problems hard and holds a commanding lead. Kimi is not a frontend slop machine, it's a generalist proto-AGI (though same can be said of Meta, congrats) h

@scaling01

my prediction for GPT-6's ECI is 166.5 (given that it's released at the end of august) so roughly ~10 ECI points higher than Kimi-K3 depending on which trendline you take this implies US frontier labs are ahead 8.2 - 10.9 months

@teortaxestex

I don't much like routing but it seems it's starting to work

@artificialanlys

Kimi K3 is second only to Fable 5 on AA-Briefcase, our agentic knowledge work benchmark, but costs more than Opus 4.8 to run while averaging nearly an hour per task Last week @Kimi_Moonshot released Kimi K3, a 2.8T parameter model that sco

@artificialanlys

AA-Briefcase measures model performance across three dimensions: binary rubric checks for ground-truth correctness, pairwise grading on analytical quality, and pairwise grading on presentation quality. The AA-Briefcase Elo is a single metri

@artificialanlys

Kimi K3’s frontier performance comes at a high Cost per Task. Its average Cost per Task of $10.57 is one of the highest recorded, below Claude Sonnet 5 (max, $14.43) and Claude Fable 5 ($22.30). Compared to Kimi K3, GPT-5.6 Sol (max) trails

@artificialanlys

Kimi K3 has an average Time per AA-Briefcase Task of 56.4 minutes. This is driven by a high number of turns, as well as higher output token use and lower speeds using the first party Kimi API. Kimi K3 uses 120k output tokens per task and 83

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive