Command Palette
Search for a command to run...

Kimi K3 Scores 85% on Vibe Code Bench to Rank Second in Open Model Coding Tests

aiai-modelingai-research-evalsai-open-models 123 posts · 48 accounts

Kimi K3 secured second place on the in-house Vibe Code Bench with an 85% score and demonstrated top-tier cybersecurity capabilities in independent evaluations, positioning the 2.8 trillion-parameter model as a leading open-weight frontier system.

The Vibe Code Bench measures end-to-end task completion, while separate internal security tests align the model's performance with major U.S. rivals. The results expand earlier coverage of the model's launch, adding verified domain-specific benchmarks to the ongoing competition in large language models.

From the sources (25 posts)

@jukan05

FT: KIMI TO UNVEIL K3 TONIGHT FT: KIMI’S MODEL IS EXPECTED TO HAVE 2–3 TRILLION PARAMETERS FT: KIMI K3 WILL BE AN OPEN-WEIGHT MODEL FT: KIMI K3 IS EXPECTED TO OUTPERFORM OPUS 4.8, BUT FALL SHORT OF FABLE 5

@jukan05

* The reporter said it could be as early as tonight. It’s not confirmed that it will be released tonight.

@synthwavedd

🚨 Per Financial Times, Moonshot intend to launch Kimi K3 as early as tonight. K3 is 2-3T total parameters - the largest Chinese model to date - while outperforming 4.8 but staying overall behind Fable in benchmarks as expected

@kimmonismus

Kimi k3 is being released tonight, via FT -2-3t parameters (Opus4.8 has about 1.5t) -1m context -Expected to exceed Opus 4.8 performance! The time when China was six months behind is over. History is presumably being made today. https://

@scaling01

RT @zijing_wu: Chinese AI start-up Moonshot to launch model challenging Anthropic’s lead * Set to release as early as tonight * 2-3T, larg…

@kimmonismus

Source:

@wallstengine

FT reports Moonshot is set to release Kimi K3 as early as tonight. K3 is expected to be China’s largest AI model to date, with 2T-3T parameters, and will be released as an open-weight model K3 is expected to outperform Claude Opus 4.8 on

@ft

RT @zijing_wu: Chinese AI start-up Moonshot to launch model challenging Anthropic’s lead * Set to release as early as tonight * 2-3T, larg…

@shanumathew93

Reminder that Opus 4.8 came out at thee end of May. So frontier leadership went from 6+ months to 1-1.5 months. Unclear how much of this is because of wide distillation and how far internal models are at American labs but the direction is

@teortaxestex

Kimi K3 clones the ad of Kimi K3

@jukan05

@zephyr_z9

RT @zijing_wu: Chinese AI start-up Moonshot to launch model challenging Anthropic’s lead * Set to release as early as tonight * 2-3T, larg…

@kimmonismus

Kimi k3 starts rolling out. Official release is imminent! Super freaking excited for the evals. Will it bear opus 4.8? Could be a real game changer, literally.

@thezachmueller

RT @xeophon: Get your cope takes before the drop: - "It isn’t at the frontier, it’s 6 months behind when you look at the unreleased Mythos…

@zephyr_z9

BRUH This is crazy Definitely Fable level at coding

@thezachmueller

RT @teortaxesTex: Preliminary, but: I think Kimi K3 could exceed Mythos with more post-training not "in theory", not "in the limit" – like,…

@teortaxestex

rare and welcome Kimi bear

@zephyr_z9

LESS GO

@ft

Chinese AI start-up Moonshot to launch model challenging Anthropic’s lead

@unusual_whales

China AI models have rapidly gained traction, per FT:

@theo

Normally I don’t comment on rumors, but if Kimi K3 actually beats out Opus 4.8 that’s nuts Also hearing Opus 5 might drop?

@synthwavedd

cute! looks like moonshot are preparing for a K3 launch today

@nrehiew_

If Kimi K3 is indeed Opus level, the main architectural ideas that would be most interesting are 1) What degree of sparsity at 2T params 2) What linear attention did they use

@synthwavedd

@theo the funniest thing is it now looks like moonshot are dropping an open-source, ~opus 5 level model before anthropic drop the actual opus 5

@scaling01

Moonshot should block Anthropic and the USG from using Kimi-K3

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive