Parameters
428.0B total / 23.0B active
MoE: total / active
Architecture
Native multimodal MoE with MiniMax Sparse Attention (MSA)
Released
02.06.2026
License
MiniMax Community License
Input Modalities
Output Modalities
Context (native)
1,000,000 tokens
Context (extended)
1,000,000 tokens
About
MiniMax-M3 (MiniMaxAI/MiniMax-M3) is a native multimodal Mixture-of-Experts model with a one-million-token context - ~428B total parameters, ~23B activated per token - handling text, image and video input, released June 2, 2026. It undergoes mixed-modality training from the very first step (early fusion), enabling deeper semantic fusion across text, image and video rather than bolting vision onto a text model.
For long-context efficiency M3 introduces MiniMax Sparse Attention (MSA), a high-performance sparse attention operator for million-token contexts that dramatically reduces attention compute and memory footprint versus GQA while preserving quality - at 1M context M3 delivers 9x prefill and 15x decode speedups over M2, cutting per-token compute to 1/20. M3 targets coding and cowork capability, achieving frontier-level performance across long-horizon agentic benchmarks. Technical report: arXiv 2606.13392.
Training Data Mixed-modality training from the first step across text, image, and video
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
SWE-bench Verified
coding_agent
|
91.86%
|
11.06.2026 |
|
SWE-bench Pro
coding_agent
|
73.75%
|
11.06.2026 |
|
Terminal Bench 2.1
coding_agent
|
72.73%
|
11.06.2026 |
|
SWE Atlas - QnA
coding_agent
|
52.37%
|
11.06.2026 |
|
NL2Repo-Bench
coding_agent
|
42.89%
|
11.06.2026 |
|
SWE Atlas - TW
coding_agent
|
41.06%
|
11.06.2026 |
|
SWE-efficiency
coding_agent
|
63.80%
|
11.06.2026 |
|
LiveSQLBench
coding_agent
|
89.74%
|
11.06.2026 |
|
CL-bench
coding_agent
|
51.00%
|
11.06.2026 |
|
VIBE-V2
coding_agent
|
79.21%
|
11.06.2026 |
|
SVG-Bench
coding_agent
|
97.52%
|
11.06.2026 |
|
PostTrainBench
coding_agent
|
81.91%
|
11.06.2026 |
|
KernelBench Hard
coding_agent
|
90.59%
|
11.06.2026 |
|
PaperBench
coding_agent
|
35.26%
|
11.06.2026 |
|
BrowseComp
general_agent
|
91.26%
|
11.06.2026 |
|
DRACO
general_agent
|
34.19%
|
11.06.2026 |
|
GDPVal
general_agent
|
87.20%
|
11.06.2026 |
|
BankerToolBench
general_agent
|
67.78%
|
11.06.2026 |
|
OfficeQA Pro
general_agent
|
55.33%
|
11.06.2026 |
|
SpreadSheetBench-v1
general_agent
|
100.00%
|
11.06.2026 |
|
YC-Bench
general_agent
|
95.45%
|
11.06.2026 |
|
MCP-Atlas
general_agent
|
84.27%
|
11.06.2026 |
|
Apex-Agents
general_agent
|
61.05%
|
11.06.2026 |
|
Claw-Eval Avg
coding_agent
|
90.09%
|
11.06.2026 |
|
OSWorld-Verified
general_agent
|
80.49%
|
11.06.2026 |
|
OmniDocBench 1.5
document_understanding
|
100.00%
|
11.06.2026 |
|
MMMU-Pro
vision_language
|
86.48%
|
11.06.2026 |
|
VideoMMMU
video_understanding
|
76.09%
|
11.06.2026 |
|
VideoMME (w sub.)
video_understanding
|
62.60%
|
11.06.2026 |
|
IMO 2025
stem_reasoning
|
100.00%
|
11.06.2026 |
|
Long-Horizon Terminal Bench
coding_agent
|
100.00%
|
16.07.2026 |
|
Humanity's Last Exam
stem_reasoning
|
66.10%
|
17.06.2026 |
|
CritPt (no tools)
stem_reasoning
|
9.78%
|
17.06.2026 |
|
HMMT Nov 25
stem_reasoning
|
38.58%
|
17.06.2026 |
|
HMMT Feb 26
stem_reasoning
|
82.77%
|
17.06.2026 |
|
GPQA Diamond
stem_reasoning
|
97.54%
|
17.06.2026 |
|
NL2Repo
coding_agent
|
48.92%
|
17.06.2026 |
|
DeepSWE 1.1
coding_agent
|
18.13%
|
17.06.2026 |
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
64.90%
|
17.06.2026 |
|
USAMO 2026
stem_reasoning
|
72.49%
|
11.06.2026 |
|
LOCA-Bench (256k)
long_context
|
49.30
|
11.06.2026 |
Model Tree, Spaces and Paper
Spaces using MiniMaxAI/MiniMax-M3 50
Collection including MiniMaxAI/MiniMax-M3
[
MiniMax-M3
Collection
7 items • Updated 27 days ago • 22
](https://huggingface.co/collections/MiniMaxAI/minimax-m3)
Paper for MiniMaxAI/MiniMax-M3
[
MiniMax Sparse Attention
Paper • 2606.13392 • Published Jun 11 • 166
](https://huggingface.co/papers/2606.13392)
Contact Us
Contact Us
Contact us at model@minimax.io.
Model size
427B params
Tensor type
BF16
·
F32
·
Model tree for MiniMaxAI/MiniMax-M3
Adapters
Finetunes
Quantizations
Local Deployment (vLLM etc.)
Local Deployment
Download the model:
hf download MiniMaxAI/MiniMax-M3 --local-dir MiniMax-M3
We recommend the following inference frameworks to serve the model:
-
SGLang - see SGLang cookbook.
-
vLLM - see vLLM recipes.
-
Transformers - see Transformers docs.
Inference Parameters
We recommend the following parameters for best performance: temperature=1.0, top_p=0.95.
How to Use
How to Use
M3 supports three reasoning modes through the thinking parameter:
enabled— Reasoning is always enabled.adaptive— M3 automatically determines when additional reasoning is beneficial.disabled— Reasoning is disabled to minimize latency and maximize throughput.
MiniMax Sparse Attention (MSA)
MiniMax Sparse Attention (MSA)
M3 is powered by MiniMax Sparse Attention (MSA), a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.

📄 Read the technical report: arXiv:2606.13392 · Hugging Face Papers
Model Overview and Highlights
MiniMaxAI/MiniMax-M3 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Page Metadata
- Downloads last month: 186,231
- Like: Like
- Show MiniMax's followers: 10,500
[
Image-Text-to-Text
](https://huggingface.co/models?pipeline_tag=image-text-to-text)[
Transformers
](https://huggingface.co/models?library=transformers)[
Safetensors
](https://huggingface.co/models?library=safetensors)[
minimax_m3_vl
](https://huggingface.co/models?other=minimax_m3_vl)[
multimodal
](https://huggingface.co/models?other=multimodal)[
Mixture of Experts
](https://huggingface.co/models?other=moe)[
agent
](https://huggingface.co/models?other=agent)[
coding
](https://huggingface.co/models?other=coding)[
video
](https://huggingface.co/models?other=video)[
conversational
](https://huggingface.co/models?other=conversational)[
custom_code
](https://huggingface.co/models?other=custom_code)[
Eval Results
](https://huggingface.co/models?other=eval-results)
Model card [Files Files and versions
xet
](https://huggingface.co/MiniMaxAI/MiniMax-M3/tree/main)[Community
29
](https://huggingface.co/MiniMaxAI/MiniMax-M3/discussions)
MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.
Highlights:
- Native Multimodality: M3 undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video.
- Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.
- Coding & Cowork Capability: M3 achieves frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.

Architecture
- Attention
- Grouped Query Attention (64:4)
- MoE
- 128 experts · top-4 per token
- Layers
- 60
- Hidden size
- 6144
- Context
- 1M tokens
- RoPE θ
- 5M
- Parameters
- 428000M
- Active params
- 23000M
Source: Hugging Face config.json · MiniMaxM3SparseForConditionalGeneration · model repo
MiniMax Sparse Attention (MSA) for million-token contexts
Training Pipeline
-
1
pretraining
Mixed-modality pretraining
Native multimodal pretraining from the very first step, enabling deeper semantic fusion across text, image, and video.
-
2
sft
Supervised fine-tuning and instruction tuning
SFT and instruction tuning with support for three reasoning modes (enabled/adaptive/disabled) via the thinking parameter.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Mixed-modality pre-training corpus (text, image, video) | pretraining | — | — |
Linked Resources
MiniMax Sparse Attention (arXiv:2606.13392)
https://arxiv.org/abs/2606.13392
MiniMax-M3 GitHub Repository
https://github.com/MiniMax-AI/MiniMax-M3
MiniMax Sparse Attention (MSA) GitHub
https://github.com/MiniMax-AI/MSA
MiniMax Agent Demo
https://agent.minimax.io/
MiniMax API Documentation
https://platform.minimax.io/
MiniMax Official Website
https://www.minimax.io/
MiniMax M3 HuggingFace Collection
https://huggingface.co/collections/MiniMaxAI/minimax-m3
Trend Analysis
24h Change
+0.3%
7d Change
+2.8%
Current
10,191
likes
+0.1%
downloads
+0.9%
downloads_all_time
+1.1%
downloads
+0.1%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 500,700 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 586,824 | daily | 01.09.2026 |
| huggingface | followers | 10,191 | daily | 01.09.2026 |
| huggingface | likes | 1,515 | daily | 01.09.2026 |
| huggingface | downloads | 196,227 | daily | 01.09.2026 |
| ollama | downloads | 500,000 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 580,652 | daily | 31.08.2026 |
| huggingface | followers | 10,159 | daily | 31.08.2026 |
| huggingface | likes | 1,514 | daily | 31.08.2026 |
| huggingface | downloads | 194,488 | daily | 31.08.2026 |
| ollama | downloads | 499,000 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 577,031 | daily | 30.08.2026 |
| huggingface | followers | 10,116 | daily | 30.08.2026 |
| huggingface | likes | 1,510 | daily | 30.08.2026 |
| huggingface | downloads | 206,031 | daily | 30.08.2026 |
| ollama | downloads | 498,200 pulls | daily | 29.08.2026 |
| huggingface | followers | 10,080 | daily | 29.08.2026 |
| huggingface | likes | 1,506 | daily | 29.08.2026 |
| huggingface | downloads | 204,823 | daily | 29.08.2026 |
| ollama | downloads | 496,900 pulls | daily | 28.08.2026 |
| huggingface | followers | 10,039 | daily | 28.08.2026 |
| huggingface | likes | 1,502 | daily | 28.08.2026 |
| huggingface | downloads | 198,826 | daily | 28.08.2026 |
| ollama | downloads | 491,200 pulls | daily | 27.08.2026 |
| huggingface | followers | 9,993 | daily | 27.08.2026 |
| huggingface | likes | 1,501 | daily | 27.08.2026 |
| huggingface | downloads | 204,245 | daily | 27.08.2026 |
| ollama | downloads | 484,200 pulls | daily | 26.08.2026 |
| huggingface | followers | 9,955 | daily | 26.08.2026 |
| huggingface | likes | 1,498 | daily | 26.08.2026 |
| huggingface | downloads | 206,380 | daily | 26.08.2026 |
| ollama | downloads | 477,300 pulls | daily | 25.08.2026 |
| huggingface | followers | 9,912 | daily | 25.08.2026 |
| huggingface | likes | 1,496 | daily | 25.08.2026 |
| huggingface | downloads | 205,085 | daily | 25.08.2026 |
| ollama | downloads | 471,000 pulls | daily | 24.08.2026 |
| huggingface | followers | 9,874 | daily | 24.08.2026 |
| huggingface | likes | 1,495 | daily | 24.08.2026 |
| huggingface | downloads | 202,878 | daily | 24.08.2026 |
| huggingface | followers | 9,831 | daily | 23.08.2026 |
| huggingface | likes | 1,494 | daily | 23.08.2026 |
| huggingface | downloads | 202,051 | daily | 23.08.2026 |
| huggingface | followers | 9,794 | daily | 22.08.2026 |
| huggingface | likes | 1,492 | daily | 22.08.2026 |
| huggingface | downloads | 202,017 | daily | 22.08.2026 |
| huggingface | followers | 9,757 | daily | 21.08.2026 |
| huggingface | likes | 1,485 | daily | 21.08.2026 |
| huggingface | downloads | 199,229 | daily | 21.08.2026 |
| huggingface | followers | 9,694 | daily | 20.08.2026 |
| huggingface | likes | 1,484 | daily | 20.08.2026 |