Moonshot AI Releases FlashKDA Kernel to Boost Kimi K3 Prefill Speed 1.72 to 2.22 Times on Nvidia H20
Moonshot AI has open-sourced FlashKDA, a high-performance attention kernel that delivers a 1.72 times to 2.22 times speedup for prompt prefilling on Nvidia H20 GPUs. The kernel functions as a drop-in replacement for existing linear attention implementations, addressing the processing delay that occurs when models read massive prompts across Kimi K3’s 1-million-token context window. The reported performance gains are specific to prefill workloads on H20 hardware and do not reflect full end-to-end inference metrics across all GPU architectures.
The company also released AgentENV, a system for training the model’s autonomous agents at scale. The framework provisions each agent with a dedicated Firecracker micro-VM that runs its own kernel, file system, memory and processes. The new tooling is available alongside the full weights and technical report for the model.
From the sources (25 posts)
@modalKimi K3 drops tomorrow. Day 0 support on Modal.
@teortaxestexmeanwhile time in China: 00:32 Monday, July 27, 2026 Strictly under 24 hours until an open frontier model
@kimmonismusQuick reminder on Kimi k3 open weights
@aaazzamRT @modal: Kimi K3 drops tomorrow. Day 0 support on Modal.
@apples_jimmyOpen weights drop in 15 hours.
@kimmonismusThey kept their promise.
@kimmonismusSource
@theoKimi is scaring the shit out of OpenAI and Anthropic. I think this is a good thing.
@togethercomputeKimi K3 lands on Together tomorrow. Available on Provisioned Throughput: reserved token-based capacity that just works: 1/ Guaranteed tok/min
@businessMoonshot AI is poised to make its Kimi K3 model available for public download, expanding its reach in the global open software community at a time of growing US concern about Chinese encroachment into the top echelons of AI development http
@cointelegraph🇨🇳 JUST IN: China's Moonshot AI is set to release its Kimi K3 model weights for free public download, widening its reach in the open-source AI race against US rivals.
@mervenoyannKimi K3 in @huggingface Inference Providers is live via @togethercompute $3/M input tokens, $15/M output tokens with 54 TPS chef's kiss
@arenaIn Frontend Code Arena, Kimi K3 (Max) by @Kimi_Moonshot is ranked #1 among open and #1 overall! It’s the #1 open-weight model in all Domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Gaming, Simulation
@mtsliveSITUATION EXPLAINED: Kimi K3's full model weights are finally out. • Moonshot released Kimi K3's weights and full technical report on July 27, about 11 days after its initial launch • 2.8 trillion parameters total, but only about 104 billi
@rohanpaul_aiKimi K3 just released their technical paper. one of the most detailed and exhaustive one. A million-token context window does not create long-horizon agency if the training system cannot preserve a task across thousands of tool calls. K
@diggMoonshot AI released Kimi K3's weights. Most users can use them for free, but a company selling K3 through an API needs a deal with @Kimi_Moonshot if it and its affiliates exceed $20M over 12 months. Products above 100M users or $20M month
@eliebakouchwill do a full deep dive on the K3 tech report later (i'm eumaxxing currently 🌴). did a first pass and it seems like an amazing tech report, laying out the blueprint of what a stable recipe at 3T scale MoE looks like. nice thread by @suchen
@tri_daoRT @togethercompute: Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch partner for @Kimi_Moonshot’s open frontier mode…
@artificialanlysKimi K3 weights have been released! Kimi K3 is now the leading open weights model at 57 in the Artificial Analysis Intelligence Index Moonshot has released the weights of their 2.6T parameter model under their 'Kimi K3 License' which we ha
@jaminballKimi K3 weights are out, and so is the pricing from 3rd parties to serve the model. Pricing from both Baseten and Fireworks is $3 / 1m input tokens, and $15 / 1m output tokens. This is $5.40 blended. The pricing is the same that Kimi of
@cursor_aiKimi K3 is now in Cursor! It scores close to the frontier on CursorBench. It's available on US-based inference thanks to our partners Fireworks, Together, and Baseten. Zero Data Retention is also supported.
@wesrothMoonshot AI has released the full Kimi K3 model weights and technical report. This is not a smaller model created specifically for open release. Kimi K3 contains 2.8 trillion total parameters, activates 104 billion parameters per token, s
@jesseproudmanRT @AskVenice: Kimi K3 by @Kimi_Moonshot is now available privately on Venice. Frontier-level capabilities, without the surveillance. http…
@togethercomputeAnother Day 0 partnership. Happy to be part of getting Kimi K3 into Cursor from launch.
@vllm_projectShoutout to @skypilot_org: awesome work shipping a full day-0 serving stack for Kimi K3 with vLLM 🙌