MiniMax-M3 Scores 55 on Artificial Analysis Intelligence Index, Would Lead Open Weights Models
Artificial Analysis scored MiniMax-M3 at 55 on its Intelligence Index, placing the model just ahead of Kimi K2.6 and MiMo-V2.5-Pro at 54, and said it would become the leading open weights model once MiniMax releases the weights, which the company has said may come in about 10 days. M3 is MiniMax's first multimodal M-series model, adding image and video input, extending the context window to 1 million tokens from 200,000 on MiniMax-M2.7, and priced at $0.30 per 1 million input tokens and $1.20 per 1 million output tokens up to 512,000 tokens of context.
On GDPval-AA, which measures real-world tasks across 44 occupations and nine major industries, M3 scored about 1,670, behind GPT-5.5 at 1,769 and Claude Opus 4.8 at 1,890, and roughly level with Claude Sonnet 4.6 at 1,676. Artificial Analysis said the model attempted only 30.9% of questions on AA-Omniscience, the lowest among current peers, yielding a 16.1% hallucination rate and 15.0% accuracy; separate earlier comparisons also showed M3 scoring above Claude Opus 4.6 on Agentic Index at 94% lower input-token cost and catching 13 planted bugs in a code audit for $0.07 versus $1.30 for the cheapest Claude Opus 4.8 run.
From the sources (6 posts)
@erikvoorheesBoth models available on the Venice API
@asvanevikOpus 4.5 felt like the breakthrough moment for agentic use of LLMs now MiniMax M3 scores higher than Opus 4.6 (!) on Agentic Index Score while costing literally 94% less per input token
@artificialanlysMiniMax-M3 scores 55 on the Artificial Analysis Intelligence Index. Once the weights are released, it will be the leading open weights model M3 is @MiniMax_AI's first multimodal M-series model, adding image and video input and a 1M token c
@artificialanlysMiniMax-M3 scores ~1670 on GDPval-AA, behind GPT-5.5 (xhigh, 1769) and level with Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort, 1676). Once the weights are released, it will be the highest-scoring open weights model on GDPval-AA. GDPva
@artificialanlysOn AA-Omniscience, MiniMax-M3 attempts only 30.9% of questions, the lowest among current peers. The abstention yields a low hallucination rate (16.1%) and accuracy (15.0%)
@minimax_aiRT @ArtificialAnlys: MiniMax-M3 scores 55 on the Artificial Analysis Intelligence Index. Once the weights are released, it will be the lead…