Kimi K3 Scores 85% on Vibe Code Bench to Rank Second in Open Model Coding Tests
Kimi K3 secured second place on the in-house Vibe Code Bench with an 85% score and demonstrated top-tier cybersecurity capabilities in independent evaluations, positioning the 2.8 trillion-parameter model as a leading open-weight frontier system.
The Vibe Code Bench measures end-to-end task completion, while separate internal security tests align the model's performance with major U.S. rivals. The results expand earlier coverage of the model's launch, adding verified domain-specific benchmarks to the ongoing competition in large language models.
From the sources (25 posts)
@jukan05FT: KIMI TO UNVEIL K3 TONIGHT FT: KIMI’S MODEL IS EXPECTED TO HAVE 2–3 TRILLION PARAMETERS FT: KIMI K3 WILL BE AN OPEN-WEIGHT MODEL FT: KIMI K3 IS EXPECTED TO OUTPERFORM OPUS 4.8, BUT FALL SHORT OF FABLE 5
@jukan05* The reporter said it could be as early as tonight. It’s not confirmed that it will be released tonight.
@synthwavedd🚨 Per Financial Times, Moonshot intend to launch Kimi K3 as early as tonight. K3 is 2-3T total parameters - the largest Chinese model to date - while outperforming 4.8 but staying overall behind Fable in benchmarks as expected
@kimmonismusKimi k3 is being released tonight, via FT -2-3t parameters (Opus4.8 has about 1.5t) -1m context -Expected to exceed Opus 4.8 performance! The time when China was six months behind is over. History is presumably being made today. https://
@scaling01RT @zijing_wu: Chinese AI start-up Moonshot to launch model challenging Anthropic’s lead * Set to release as early as tonight * 2-3T, larg…
@kimmonismusSource:
@wallstengineFT reports Moonshot is set to release Kimi K3 as early as tonight. K3 is expected to be China’s largest AI model to date, with 2T-3T parameters, and will be released as an open-weight model K3 is expected to outperform Claude Opus 4.8 on
@ftRT @zijing_wu: Chinese AI start-up Moonshot to launch model challenging Anthropic’s lead * Set to release as early as tonight * 2-3T, larg…
@shanumathew93Reminder that Opus 4.8 came out at thee end of May. So frontier leadership went from 6+ months to 1-1.5 months. Unclear how much of this is because of wide distillation and how far internal models are at American labs but the direction is
@teortaxestexKimi K3 clones the ad of Kimi K3
@jukan05
@zephyr_z9RT @zijing_wu: Chinese AI start-up Moonshot to launch model challenging Anthropic’s lead * Set to release as early as tonight * 2-3T, larg…
@kimmonismusKimi k3 starts rolling out. Official release is imminent! Super freaking excited for the evals. Will it bear opus 4.8? Could be a real game changer, literally.
@thezachmuellerRT @xeophon: Get your cope takes before the drop: - "It isn’t at the frontier, it’s 6 months behind when you look at the unreleased Mythos…
@zephyr_z9BRUH This is crazy Definitely Fable level at coding
@thezachmuellerRT @teortaxesTex: Preliminary, but: I think Kimi K3 could exceed Mythos with more post-training not "in theory", not "in the limit" – like,…
@teortaxestexrare and welcome Kimi bear
@zephyr_z9LESS GO
@ftChinese AI start-up Moonshot to launch model challenging Anthropic’s lead
@unusual_whalesChina AI models have rapidly gained traction, per FT:
@theoNormally I don’t comment on rumors, but if Kimi K3 actually beats out Opus 4.8 that’s nuts Also hearing Opus 5 might drop?
@synthwaveddcute! looks like moonshot are preparing for a K3 launch today
@nrehiew_If Kimi K3 is indeed Opus level, the main architectural ideas that would be most interesting are 1) What degree of sparsity at 2T params 2) What linear attention did they use
@synthwavedd@theo the funniest thing is it now looks like moonshot are dropping an open-source, ~opus 5 level model before anthropic drop the actual opus 5
@scaling01Moonshot should block Anthropic and the USG from using Kimi-K3