Moonshot Open-Sources FlashKDA, Highlights Kimi K2.6 Coding Benchmarks
Moonshot said it is open-sourcing FlashKDA, a CUTLASS-based implementation of Kimi Delta Attention kernels that can serve as a drop-in backend for flash-linear-attention. The company said FlashKDA delivers 1.72x to 2.22x faster prefill than the flash-linear-attention baseline on H20.
In separate demonstrations, Moonshot said Kimi K2.6 sustained 12- to 13-hour coding runs with more than 4,000 tool calls, reworking the open-source exchange-core matching engine and optimizing local Zig-based inference for Qwen3.5-0.8B on Mac. It reported throughput gains from 0.43 to 1.24 MT/s and from 1.23 to 2.86 MT/s on exchange-core, and from about 15 to about 193 tokens per second in the local inference test.
From the sources (3 posts)
@kimi_moonshotWe're open-sourcing FlashKDA — our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. Achieves 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash
@kimi_moonshotKimi K2.6 autonomously overhauled exchange-core, an 8-year-old open-source financial matching engine. Over a 13-hour execution, the model iterated through 12 optimization strategies, initiating over 1,000 tool calls to precisely modify mo
@kimi_moonshotKimi K2.6 demonstrates strong long-horizon coding in complex engineering tasks: Kimi K2.6 successfully downloaded and deployed the Qwen3.5-0.8B model locally on a Mac. By implementing and optimizing model inference in Zig—a highly niche pr