MiniMax Releases Open-Weight M3, a 428B Multimodal Model With 1M-Token Context
MiniMax released the open-weight M3 model on Hugging Face, describing it as a native multimodal system for coding and agentic workloads with a 1 million-token context window. M3 has about 428 billion total parameters, with roughly 23 billion activated, and uses MiniMax Sparse Attention; the company said it scored 59.0% on SWE-Bench Pro and 66.0% on Terminal Bench 2.1 while supporting text, image and video from the outset.
The launch arrived with broad deployment support. vLLM said M3 has day-zero support verified on Nvidia and AMD hardware, describing the sparse-attention design as scoring 128-token KV blocks and attending only the top blocks to make 1 million-token serving practical. Together AI also made the model available in the cloud, claiming up to 125% higher throughput, while Unsloth said M3 can run locally; MiniMax said a technical report will follow in about 10 days.
From the sources (25 posts)
@minimax_aiRT @RyanLeeMiniMax: Hey everyone — our high-performance MSA kernel library is now open-source. The M3 weights are expected to drop this Fri…
@minimax_aiRT @RyanLeeMiniMax: Hey everyone — our high-performance MSA kernel library is now open-source. The M3 weights are expected to drop this Fri…
@minimax_aiWeights on Friday 🫶
@minimax_aiMiniMax M3, Open-Weight, Now On Hugging Face Weights: MiniMax Sparse Attention:
@minimax_aiRT @lmsysorg: 🎉 SGLang has Day-0 support for MiniMax-M3 from @MiniMax_AI, a native-multimodal MoE reasoning model of ~428B total params (~2…
@minimax_aiRT @RyanLeeMiniMax: On the M3 license — thanks for the feedback on M2.7 You told us prior-approval-for-any-commercial-use was too much. We…
@minimax_aiMiniMax M3, Open-Weight, Now On Hugging Face , with only ~428B parameters and ~23B activated parameters Weights: MiniMax Sparse Attention:
@yacinemtbRT @MiniMax_AI: MiniMax M3, Open-Weight, Now On Hugging Face , with only ~428B parameters and ~23B activated parameters Weights: https://t…
@lmsysorg🎉 SGLang has Day-0 support for MiniMax-M3 from @MiniMax_AI, a native-multimodal MoE reasoning model of ~428B total params (~23B active), 60 layers, 1M context across text, image & video. ✅ Native multimodality: text-image-video fusion from
@adinayakupMiniMax-M3 just dropped on @huggingface ✨ 428B / 23B active ✨ 1M context ✨ MiniMax Sparse Attention (MSA) And it’s not just weights! Day-one full release: - paper - kernel - Transformers support Love how this was released❤️ @MiniMax_AI
@vllm_project🎉 Congrats to @MiniMax_AI on releasing MiniMax M3! Frontier coding and agentic capabilities, native image and video input, computer use, and a 1M-token context window, all in a single open model. At the heart of M3 is MSA, a new sparse att
@vllm_projectDay-0 goes beyond inference: NeMo RL from @NVIDIAAI also supports MiniMax M3 on day 0, with vLLM powering rollout generation. 💡 A reference GRPO recipe is ready, so you can start post-training M3 for your own agentic workflows right away.
@nvidiaaiCongrats to the @MiniMax_AI team on the release of MiniMax M3, a long-context multimodal model for text, image, and video reasoning. 🙌 Try it today with our free GPU-accelerated endpoint on Details:
@zephyr_z9with only ~428B parameters and ~23B activated parameters Smaller than I expected
@minimax_aiM3 open weight just dropped and it's live on @Modular cloud on day zero with up to a 1M-context and MSA architecture kernel-to-cloud optimization is exactly what M3 needs glad to have @Modular with us from the start
@_akhaliqRT @novita_labs: 🤗 MiniMax M3 from @MiniMax_AI is now live on @huggingface — supported by Novita. Open weights. ~428B total parameters. ~2…
@minimax_aiM3 is now live on @parasail_io 🚀
@clattner_llvmM3 from @MiniMax_AI is now live on Modular Cloud, day zero. Open weights, 1M-token context, multimodal. MSA is a new attn architecture and getting its performance benefits requires whole-stack optimization. That's what we built Modular fo
@unslothaiMiniMax M3 can now be run locally!🔥 MiniMax-M3 is a new 428B (23B active) open model with 1M context that performs on par with Gemini 3.1 Pro. Run Dynamic 2-bit GGUF on 138GB RAM/VRAM or 3-bit on 165GB. GGUF: Guid
@minimax_aiRun M3 locally today with @UnslothAI
@far__elRT @MiniMax_AI: Introducing MiniMax M3: The First Open-Weights Model to Combine Three Frontier Capabilities - Coding & Agentic Frontier: 5…
@far__elRT @MiniMax_AI: M3 would never 🙂↔️ As a matter of fact, the weights are now open, too.
@danielhanchenRT @UnslothAI: MiniMax M3 can now be run locally!🔥 MiniMax-M3 is a new 428B (23B active) open model with 1M context that performs on par w…
@togethercomputeMiniMax-M3 from @MiniMax_AI is now available on Together AI. It’s an open-weight native multimodal model with 1M context, MiniMax Sparse Attention, and thinking / non-thinking modes. Together AI is MiniMax’s preferred cloud partner, with
@minimax_aithe kernels are doing the lord's work today, day-0 on @vllm_project, verified on nvidia and amd. go read the writeup 👇