Command Palette
Search for a command to run...

vLLM Launches 2.8T-Parameter Kimi K3 Deployment on Nvidia and AMD Hardware

aiai-infrastructureai-inference-platformsai-compute-chipsai-modelingai-open-models 55 posts · 34 accounts

vLLM launches day-zero deployment of Moonshot AI’s Kimi K3 model on Nvidia and AMD computing hardware. The inference platform runs the 2.8 trillion parameter model using its engine on Nvidia’s Grace Hopper and Blackwell architectures, alongside AMD Instinct GPUs, enabling production-scale serving immediately upon the model’s public release.

The launch expands initial infrastructure options for the open-weight model, which Moonshot AI released alongside a technical report this week. In addition to bare-metal hardware support, inference endpoints on Modal, Baseten and DigitalOcean will be available to developers starting July 29. vLLM is also rolling out broader performance tuning for AMD systems in upcoming updates.

From the sources (25 posts)

@apples_jimmy

Open weights drop in 15 hours.

@kimmonismus

They kept their promise.

@kimmonismus

Source

@theo

Kimi is scaring the shit out of OpenAI and Anthropic. I think this is a good thing.

@togethercompute

Kimi K3 lands on Together tomorrow. Available on Provisioned Throughput: reserved token-based capacity that just works: 1/ Guaranteed tok/min

@business

Moonshot AI is poised to make its Kimi K3 model available for public download, expanding its reach in the global open software community at a time of growing US concern about Chinese encroachment into the top echelons of AI development http

@cointelegraph

🇨🇳 JUST IN: China's Moonshot AI is set to release its Kimi K3 model weights for free public download, widening its reach in the open-source AI race against US rivals.

@mervenoyann

Kimi K3 in @huggingface Inference Providers is live via @togethercompute $3/M input tokens, $15/M output tokens with 54 TPS chef's kiss

@arena

In Frontend Code Arena, Kimi K3 (Max) by @Kimi_Moonshot is ranked #1 among open and #1 overall! It’s the #1 open-weight model in all Domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Gaming, Simulation

@mtslive

SITUATION EXPLAINED: Kimi K3's full model weights are finally out. • Moonshot released Kimi K3's weights and full technical report on July 27, about 11 days after its initial launch • 2.8 trillion parameters total, but only about 104 billi

@rohanpaul_ai

Kimi K3 just released their technical paper. one of the most detailed and exhaustive one. A million-token context window does not create long-horizon agency if the training system cannot preserve a task across thousands of tool calls. K

@digg

Moonshot AI released Kimi K3's weights. Most users can use them for free, but a company selling K3 through an API needs a deal with @Kimi_Moonshot if it and its affiliates exceed $20M over 12 months. Products above 100M users or $20M month

@eliebakouch

will do a full deep dive on the K3 tech report later (i'm eumaxxing currently 🌴). did a first pass and it seems like an amazing tech report, laying out the blueprint of what a stable recipe at 3T scale MoE looks like. nice thread by @suchen

@tri_dao

RT @togethercompute: Kimi K3 is now live on Together AI. We’re proud to be a Day 0 launch partner for @Kimi_Moonshot’s open frontier mode…

@artificialanlys

Kimi K3 weights have been released! Kimi K3 is now the leading open weights model at 57 in the Artificial Analysis Intelligence Index Moonshot has released the weights of their 2.6T parameter model under their 'Kimi K3 License' which we ha

@jaminball

Kimi K3 weights are out, and so is the pricing from 3rd parties to serve the model. Pricing from both Baseten and Fireworks is $3 / 1m input tokens, and $15 / 1m output tokens. This is $5.40 blended. The pricing is the same that Kimi of

@cursor_ai

Kimi K3 is now in Cursor! It scores close to the frontier on CursorBench. It's available on US-based inference thanks to our partners Fireworks, Together, and Baseten. Zero Data Retention is also supported.

@wesroth

Moonshot AI has released the full Kimi K3 model weights and technical report. This is not a smaller model created specifically for open release. Kimi K3 contains 2.8 trillion total parameters, activates 104 billion parameters per token, s

@jesseproudman

RT @AskVenice: Kimi K3 by @Kimi_Moonshot is now available privately on Venice. Frontier-level capabilities, without the surveillance. http…

@togethercompute

Another Day 0 partnership. Happy to be part of getting Kimi K3 into Cursor from launch.

@vllm_project

Shoutout to @skypilot_org: awesome work shipping a full day-0 serving stack for Kimi K3 with vLLM 🙌

@art_zucker

RT @Kimi_Moonshot: Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with n…

@_akhaliq

RT @jeffboudier: Kimi K3 is available day-0 for on-premise deployment on @Dell Host the 2.8T frontier open intelligence in your own serve…

@baseten

RT @feilsystem: We shipped the worlds fastest tokenizer for KimiK3. Basetenkenizer now available on day 0 on Baseten and on pypi. https://t…

@fs0c131y

RT @Kimi_Moonshot: Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with n…

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive