MiniMax-M3
MiniMax's natively multimodal open-weight flagship uniting frontier reasoning and agentic capability in one model, with sparse attention for ultra-long context
MiniMax-M3
MiniMax • June 2026
Training Data
Up to early 2026
MiniMax-M3
June 2026
Parameters
428 billion (23B active)
Training Method
Mixture of Experts with MiniMax Sparse Attention (MSA)
Context Window
1,000,000 tokens
Knowledge Cutoff
Not disclosed
Key Features
Native Multimodal (Image + Video) • 1M Context via Sparse Attention • Computer Use • Open Weights
Capabilities
Coding: Outstanding
Agent Tasks: Outstanding
Multimodal: Excellent
What's New in This Version
First open-weights model to unite reasoning and agent capabilities in one release; MSA delivers up to 15.6x faster long-context response and scores 59.0% on SWE-Bench Pro, a major leap over the text-only M2.7
MiniMax's natively multimodal open-weight flagship uniting frontier reasoning and agentic capability in one model, with sparse attention for ultra-long context
What's New in This Version
First open-weights model to unite reasoning and agent capabilities in one release; MSA delivers up to 15.6x faster long-context response and scores 59.0% on SWE-Bench Pro, a major leap over the text-only M2.7
Technical Specifications
Key Features
Capabilities
Other MiniMax Models
Explore more models from MiniMax
MiniMax-M2.7
MiniMax's self-evolving agent model pioneering recursive self-improvement with frontier agentic coding performance at a fraction of competitor cost
MiniMax-M2.5
MiniMax's flagship model matching frontier performance at 1/20th the cost with 80.2% SWE-bench Verified
MiniMax-M2.5-Lightning
Ultra-fast variant of M2.5 generating 100 tokens per second at $1/hour continuous operation