MiniMax-M3

MiniMaxAI

Parameters

428.0B total / 23.0B active

MoE: total / active

Architecture

Native multimodal MoE with MiniMax Sparse Attention (MSA)

Released

02.06.2026

License

MiniMax Community License

Open Weights Commercial Use Multimodal BF16, F32 MiniMax en zh multilingual

Input Modalities

text image video

Output Modalities

text

Context (native)

1,000,000 tokens

Context (extended)

1,000,000 tokens

Openness Index Score 70.0/100

About

MiniMax-M3 (MiniMaxAI/MiniMax-M3) is a native multimodal Mixture-of-Experts model with a one-million-token context - ~428B total parameters, ~23B activated per token - handling text, image and video input, released June 2, 2026. It undergoes mixed-modality training from the very first step (early fusion), enabling deeper semantic fusion across text, image and video rather than bolting vision onto a text model.

For long-context efficiency M3 introduces MiniMax Sparse Attention (MSA), a high-performance sparse attention operator for million-token contexts that dramatically reduces attention compute and memory footprint versus GQA while preserving quality - at 1M context M3 delivers 9x prefill and 15x decode speedups over M2, cutting per-token compute to 1/20. M3 targets coding and cowork capability, achieving frontier-level performance across long-horizon agentic benchmarks. Technical report: arXiv 2606.13392.

Training Data Mixed-modality training from the first step across text, image, and video

Benchmark Scores

Benchmark Score Date
SWE-bench Verified
coding_agent
91.86%
11.06.2026
SWE-bench Pro
coding_agent
73.75%
11.06.2026
Terminal Bench 2.1
coding_agent
72.73%
11.06.2026
SWE Atlas - QnA
coding_agent
52.37%
11.06.2026
NL2Repo-Bench
coding_agent
42.89%
11.06.2026
SWE Atlas - TW
coding_agent
41.06%
11.06.2026
SWE-efficiency
coding_agent
63.80%
11.06.2026
LiveSQLBench
coding_agent
89.74%
11.06.2026
CL-bench
coding_agent
51.00%
11.06.2026
VIBE-V2
coding_agent
79.21%
11.06.2026
SVG-Bench
coding_agent
97.52%
11.06.2026
PostTrainBench
coding_agent
81.91%
11.06.2026
KernelBench Hard
coding_agent
90.59%
11.06.2026
PaperBench
coding_agent
35.26%
11.06.2026
BrowseComp
general_agent
91.26%
11.06.2026
DRACO
general_agent
34.19%
11.06.2026
GDPVal
general_agent
87.20%
11.06.2026
BankerToolBench
general_agent
67.78%
11.06.2026
OfficeQA Pro
general_agent
55.33%
11.06.2026
SpreadSheetBench-v1
general_agent
100.00%
11.06.2026
YC-Bench
general_agent
95.45%
11.06.2026
MCP-Atlas
general_agent
84.27%
11.06.2026
Apex-Agents
general_agent
61.05%
11.06.2026
Claw-Eval Avg
coding_agent
90.09%
11.06.2026
OSWorld-Verified
general_agent
80.49%
11.06.2026
OmniDocBench 1.5
document_understanding
100.00%
11.06.2026
MMMU-Pro
vision_language
86.48%
11.06.2026
VideoMMMU
video_understanding
76.09%
11.06.2026
VideoMME (w sub.)
video_understanding
62.60%
11.06.2026
IMO 2025
stem_reasoning
100.00%
11.06.2026
Long-Horizon Terminal Bench
coding_agent
100.00%
16.07.2026
Humanity's Last Exam
stem_reasoning
66.10%
17.06.2026
CritPt (no tools)
stem_reasoning
9.78%
17.06.2026
HMMT Nov 25
stem_reasoning
38.58%
17.06.2026
HMMT Feb 26
stem_reasoning
82.77%
17.06.2026
GPQA Diamond
stem_reasoning
97.54%
17.06.2026
NL2Repo
coding_agent
48.92%
17.06.2026
DeepSWE 1.1
coding_agent
18.13%
17.06.2026
Terminal-Bench 2.1 (Terminus-2)
coding_agent
64.90%
17.06.2026
USAMO 2026
stem_reasoning
72.49%
11.06.2026
LOCA-Bench (256k)
long_context
49.30
11.06.2026

Model Tree, Spaces and Paper

Spaces using MiniMaxAI/MiniMax-M3 50

Collection including MiniMaxAI/MiniMax-M3

[

MiniMax-M3

Collection

7 items • Updated 27 days ago • 22

](https://huggingface.co/collections/MiniMaxAI/minimax-m3)

Paper for MiniMaxAI/MiniMax-M3

[

MiniMax Sparse Attention

Paper • 2606.13392 • Published Jun 11 • 166

](https://huggingface.co/papers/2606.13392)

Contact Us

Contact Us

Contact us at model@minimax.io.

Safetensors

Model size

427B params

Tensor type

BF16

·

F32

·

Model tree for MiniMaxAI/MiniMax-M3

Adapters

1 model

Finetunes

13 models

Quantizations

59 models

Local Deployment (vLLM etc.)

Local Deployment

Download the model:

hf download MiniMaxAI/MiniMax-M3 --local-dir MiniMax-M3

We recommend the following inference frameworks to serve the model:

Inference Parameters

We recommend the following parameters for best performance: temperature=1.0, top_p=0.95.

How to Use

How to Use

M3 supports three reasoning modes through the thinking parameter:

  • enabled — Reasoning is always enabled.
  • adaptive — M3 automatically determines when additional reasoning is beneficial.
  • disabled — Reasoning is disabled to minimize latency and maximize throughput.

MiniMax Sparse Attention (MSA)

MiniMax Sparse Attention (MSA)

M3 is powered by MiniMax Sparse Attention (MSA), a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.

GQA vs MSA Efficiency Comparison

📄 Read the technical report: arXiv:2606.13392 · Hugging Face Papers

Model Overview and Highlights

MiniMaxAI/MiniMax-M3 · Hugging Face

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Page Metadata

  • Downloads last month: 186,231
  • Like: Like
  • Show MiniMax's followers: 10,500

Hugging Face

[

Image-Text-to-Text

](https://huggingface.co/models?pipeline_tag=image-text-to-text)[

Transformers

](https://huggingface.co/models?library=transformers)[

Safetensors

](https://huggingface.co/models?library=safetensors)[

minimax_m3_vl

](https://huggingface.co/models?other=minimax_m3_vl)[

multimodal

](https://huggingface.co/models?other=multimodal)[

Mixture of Experts

](https://huggingface.co/models?other=moe)[

agent

](https://huggingface.co/models?other=agent)[

coding

](https://huggingface.co/models?other=coding)[

video

](https://huggingface.co/models?other=video)[

conversational

](https://huggingface.co/models?other=conversational)[

custom_code

](https://huggingface.co/models?other=custom_code)[

Eval Results

](https://huggingface.co/models?other=eval-results)

Model card [Files Files and versions

xet

](https://huggingface.co/MiniMaxAI/MiniMax-M3/tree/main)[Community

29

](https://huggingface.co/MiniMaxAI/MiniMax-M3/discussions)



ModelScope MiniMax AI

MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.

Highlights:

  • Native Multimodality: M3 undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video.
  • Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.
  • Coding & Cowork Capability: M3 achieves frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.

![](https://huggingface.co/MiniMaxAI/MiniMax-M3/resolve/main/figures/ben

(section continues in the model card)

Architecture

Decoder Block input Embedding vocab 200K · d 6144 Full Attention GQA 64:4 · dₕ 128 ×60 MoE FFN 128 experts · top-4 · +1 shared · dᴻ 3072 MTP Head ×1 speculative layer Final Norm LM Head vocab 200K output
Attention
Grouped Query Attention (64:4)
MoE
128 experts · top-4 per token
Layers
60
Hidden size
6144
Context
1M tokens
RoPE θ
5M
Parameters
428000M
Active params
23000M

Source: Hugging Face config.json · MiniMaxM3SparseForConditionalGeneration · model repo

Type: Mixture of Experts (MoE)
Attention: MiniMax Sparse Attention (MSA)
Decoder: Transformer
MoE: yes (? experts)
Routing: Sparse attention with GQA baseline replacement
Total parameters 428B
Active parameters 23B
Context length 1M tokens native multimodal
Precision BF16 / F32
Compute Reduction 1/20 per-token at 1M context
Speedup Vs M2 Decode 15x
Speedup Vs M2 Prefill 9x
Key Innovation

MiniMax Sparse Attention (MSA) for million-token contexts

Training Pipeline

  1. 1
    pretraining

    Mixed-modality pretraining

    Native multimodal pretraining from the very first step, enabling deeper semantic fusion across text, image, and video.

  2. 2
    sft

    Supervised fine-tuning and instruction tuning

    SFT and instruction tuning with support for three reasoning modes (enabled/adaptive/disabled) via the thinking parameter.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Mixed-modality pre-training corpus (text, image, video) pretraining — —

Trend Analysis

24h Change

+0.3%

7d Change

+2.8%

Current

10,191

huggingface

likes

+0.1%

huggingface

downloads

+0.9%

huggingface

downloads_all_time

+1.1%

ollama

downloads

+0.1%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 500,700 pulls daily 01.09.2026
huggingface downloads_all_time 586,824 daily 01.09.2026
huggingface followers 10,191 daily 01.09.2026
huggingface likes 1,515 daily 01.09.2026
huggingface downloads 196,227 daily 01.09.2026
ollama downloads 500,000 pulls daily 31.08.2026
huggingface downloads_all_time 580,652 daily 31.08.2026
huggingface followers 10,159 daily 31.08.2026
huggingface likes 1,514 daily 31.08.2026
huggingface downloads 194,488 daily 31.08.2026
ollama downloads 499,000 pulls daily 30.08.2026
huggingface downloads_all_time 577,031 daily 30.08.2026
huggingface followers 10,116 daily 30.08.2026
huggingface likes 1,510 daily 30.08.2026
huggingface downloads 206,031 daily 30.08.2026
ollama downloads 498,200 pulls daily 29.08.2026
huggingface followers 10,080 daily 29.08.2026
huggingface likes 1,506 daily 29.08.2026
huggingface downloads 204,823 daily 29.08.2026
ollama downloads 496,900 pulls daily 28.08.2026
huggingface followers 10,039 daily 28.08.2026
huggingface likes 1,502 daily 28.08.2026
huggingface downloads 198,826 daily 28.08.2026
ollama downloads 491,200 pulls daily 27.08.2026
huggingface followers 9,993 daily 27.08.2026
huggingface likes 1,501 daily 27.08.2026
huggingface downloads 204,245 daily 27.08.2026
ollama downloads 484,200 pulls daily 26.08.2026
huggingface followers 9,955 daily 26.08.2026
huggingface likes 1,498 daily 26.08.2026
huggingface downloads 206,380 daily 26.08.2026
ollama downloads 477,300 pulls daily 25.08.2026
huggingface followers 9,912 daily 25.08.2026
huggingface likes 1,496 daily 25.08.2026
huggingface downloads 205,085 daily 25.08.2026
ollama downloads 471,000 pulls daily 24.08.2026
huggingface followers 9,874 daily 24.08.2026
huggingface likes 1,495 daily 24.08.2026
huggingface downloads 202,878 daily 24.08.2026
huggingface followers 9,831 daily 23.08.2026
huggingface likes 1,494 daily 23.08.2026
huggingface downloads 202,051 daily 23.08.2026
huggingface followers 9,794 daily 22.08.2026
huggingface likes 1,492 daily 22.08.2026
huggingface downloads 202,017 daily 22.08.2026
huggingface followers 9,757 daily 21.08.2026
huggingface likes 1,485 daily 21.08.2026
huggingface downloads 199,229 daily 21.08.2026
huggingface followers 9,694 daily 20.08.2026
huggingface likes 1,484 daily 20.08.2026

View full metric history →

Related Models