Moonshot’s Kimi K3 Fixes 15 Security Bugs OpenAI and Anthropic Systems Refused to Patch
Moonshot AI’s Kimi K3 patched 15 critical security vulnerabilities that OpenAI’s Codex and Anthropic’s Claude Fable 5 refused to address due to built-in safety filters. The open-weight model’s execution in a real-world cybersecurity task triggered fresh comparisons between American guardrailed systems and Chinese frontier AI.
Separate testing on a dedicated cybersecurity benchmark ranked Kimi K3 as the strongest open-source variant for the discipline. The results provide immediate training data for reasoning models and have intensified market scrutiny over whether U.S. AI safety protocols are hindering competitive delivery in high-value technical work.
From the sources (25 posts)
@crystalsssupRT @AfterQuery: Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5. An open weight model now outperforms al…
@fs0c131yRT @Kimi_Moonshot: Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi…
@nandodfRT @SemiAnalysis_: CHINA’S KIMI K3 HAS SURPASSED ALL AMERICAN MODELS IN FRONT-END CODING WHILE BEING SMALLER THAN MOST CLOSED-SOURCE FRONTI…
@semianalysis_A year ago, the big three was OpenAI, Anthropic, and Google. Things have changed. Moonshot's Kimi K3 sits above Gemini on every composite benchmark, and it's open source in 10 days. New episode: what K3 reveals about frontier margins, mod
@businessThe surprising gains made by Moonshot's Kimi have stunned some AI watchers and touched off concerns about whether the immense spending commitments from Silicon Valley will pay off.
@selkis_2028RT @cgtwts: Moonshot AI CEO Yang Zhilin just gave one of the clearest masterclasses on how frontier AI models are actually built. https://t…
@lfg_capRT @AymericRoucher: Kimi-K3 is not really better than Fable, and it's still lagging months behind the closed-source frontier. But it IS co…
@wsjThe rise of products such as Moonshot’s Kimi K3 raises concern about the tech boom’s staying power.
@zookoRT @robustdragon: I don’t care whether Kimi wins, DeepSeek wins, OpenAI wins, or Anthropic wins. I care that no one wins the monopoly. Th…
@scaling01RT @voxelbench: Kimi K3 ranks 3rd on VoxelBench just 100+ Elo points behind Fable! and a huge uplift from K2.6 (28th)
@kimi_moonshotRT @voxelbench: Kimi K3 ranks 3rd on VoxelBench just 100+ Elo points behind Fable! and a huge uplift from K2.6 (28th)
@zookoRT @lemire: How should we think about the impact of better AI models? Yesterday, a Chinese firm announced a much better model, one that can…
@novogratzWe need to rethink our stance to being so anti-immigrant. While correcting open border policies was necessary we are way over correcting. America needs 2mm immigrants a year. This year we are probably zero. We need both labor and highl
@zookoRT @deanwball: Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation…
@kimi_moonshotFeeling all the love for Kimi K3 already. Here are some of the amazing things people have been building with it. Enjoy K3.
@dmjk001彭博: 月之暗面 Kimi 颠覆美国领先中国的传统认知 月之暗面发布 Kimi K3 模型,其综合能力超越除 Anthropic 的 Claude Fable 5 和 OpenAI 的 GPT-5.6 之外的所有竞品。 Kimi K3 的发布令部分 AI 观察人士与投资者震惊,加剧了市场对硅谷巨额投入能否获得回报的担忧,并动摇了美国技术领先地位的信心。 月之暗面的突破可能使美国监管机构保护新模型的努力复杂化,并凭借 K3 在编程等高价值任务上的卓越表现,以高端产品定位对 Op
@garrytanRT @datacurve: Kimi K3 debuts at #3 on DeepSWE. It's the first open-weights model that delivers frontier-level performance, achieving resu…
@kimi_moonshotRT @togethercompute: We analyzed Kimi K3 vs. Claude Fable 5 for software engineering tasks using DeepSWE. Kimi K3 gets you the same perfor…
@clashreportChina's Moonshot AI has reset the global tech balance by launching world's largest open-weight system. ▪️Features 2.8 trillion parameters, 1M token window ▪️Defeated top US systems in front-end coding tests ▪️Erased US frontier lead at 40%
@nic_carterLooks like the AI safetyist decel pause AI crowd are the biggest losers in this whole Kimi thing. And that swells my heart with joy
@teortaxestexThis is a pretty compelling demonstration that Kimi K3 is both generally intelligent (on some level at least) and generally knowledgeable, and in fact knowledgeable in exactly the dual-use domain that makes Dario lose sleep.
@rauchgBased on internal evals: ▪️ Kimi K3 is top-tier at cybersecurity There is chatter on X that Moonshot benchmark-overfit. These are stealth evals. Model has raw IQ. ▪️ Sol is a leap ahead in cyber capability At a significantly higher cost,
@andersonbcdefgRT @deredleritt3r: Added to prinzbench: Kimi K3. K3 is by far the best open-source model I have tested to date. Its overall score (47/99)…
@teortaxestexKimi is very decent on cybersecurity. GPT 5.5 level One more benchmark where it shows real strength.
@andrewcurran_RT @cramforce: We ran Kimi K3 on a private cybersecurity benchmark. TL;DR: Kimi K3 is the workhorse for cyber security tasks at great rec…