DeepSeek-V4-Flash

DeepSeek

Parameters

291.0B total / 13.0B active

MoE: total / active

Architecture

MoE

Released

26.06.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16, I64, F32, F8_E4M3, I8 DeepSeek en zh

Input Modalities

text

Output Modalities

text

Context (native)

1,000,000 tokens

Context (extended)

1,000,000 tokens

Openness Index Score 100.0/100

About

DeepSeek-V4-Flash (deepseek-ai/DeepSeek-V4-Flash) is the preview release of the smaller DeepSeek-V4 MoE model - 284B total parameters with 13B activated per token (card figure; DB stores 291B total) - supporting a one-million-token context, released June 26, 2026 under the MIT License. It supports three reasoning modes: Non-think, Think High, Think Max.

The DeepSeek-V4 architecture combines Compressed Sparse Attention (CSA) with Heavily Compressed Attention (HCA) in a hybrid stack that dramatically improves long-context efficiency - at 1M context the series needs only 27% of single-token inference FLOPs and 10% of the KV cache compared with DeepSeek-V3.2 - and strengthens residual connections with Manifold-Constrained Hyper-Connections (mHC) for stable cross-layer signal propagation. Training uses the Muon optimizer for faster convergence.

Pre-trained on 32T+ diverse high-quality tokens, its post-training runs a two-stage paradigm: independent cultivation of domain-specific experts via SFT and RL with GRPO, then unified model consolidation via on-policy distillation into a single model. DeepSeek-V4-Flash-Max (maximum reasoning effort) achieves reasoning performance comparable to the Pro version given a larger thinking budget, though its smaller scale places it slightly behind on pure knowledge and the most complex agentic workflows.

Training Data Preview version of DeepSeek-V4-Flash. Pre-trained on 32T+ diverse tokens. Post-training: two-stage paradigm with independent domain-specific expert cultivation (SFT + RL with GRPO), then unified model consolidation via on-policy distillation. Supports three reasoning modes: Non-think, Think High, Think Max.

Benchmark Scores

Benchmark Score Date
BrowseComp
general_agent
79.57%
26.06.2026
HLE (with tools)
stem_reasoning
58.90%
26.06.2026
Vals.ai Financial Agent 1.1 (without web search)
general_agent
71.00%
04.06.2026
Terminal-Bench 2.1 (Terminus-2)
coding_agent
48.97%
04.06.2026
Vals.ai Financial Agent 1.1 (with web search)
general_agent
81.36%
04.06.2026
CritPt (no tools)
stem_reasoning
31.55%
04.06.2026
GDPVal
general_agent
33.84%
04.06.2026
LiveCodeBench v6
stem_reasoning
97.41%
26.06.2026
IMOAnswerBench
stem_reasoning
93.47%
26.06.2026
SWE-bench Multilingual
coding_agent
80.83%
26.06.2026
OmniScience Accuracy
knowledge
73.76%
04.06.2026
ProfBench (Search)
general_agent
79.14%
04.06.2026
IMOAnswerBench (with tools)
stem_reasoning
77.92%
04.06.2026
OmniScience Non-Hallucination
knowledge
2.80
04.06.2026
PinchBench
general_agent
100.00%
04.06.2026
TauBench V3 Airline
general_agent
52.38%
04.06.2026
IFBench (prompt loose)
instruction_following
100.00%
04.06.2026
Apex-Shortlist (with tools)
stem_reasoning
86.99%
04.06.2026
Multi-Challenge
instruction_following
97.22%
04.06.2026
TauBench V3 Retail
general_agent
100.00%
04.06.2026
AA-LCR
long_context
78.38%
04.06.2026
RULER (1M)
long_context
87.70
04.06.2026
TauBench V3 Telecom
general_agent
100.00%
04.06.2026
TauBench V3 Banking
general_agent
67.48%
04.06.2026
SciCode (subtask)
stem_reasoning
48.06%
04.06.2026
Longbench v2 (≤ 1M)
long_context
57.00
04.06.2026
MMLU-ProX
multilingual
86.45%
04.06.2026
WMT24++ (en→xx)
multilingual
98.20%
04.06.2026
TauBench V3 Average
general_agent
100.00%
04.06.2026
Terminal Bench 2.1
coding_agent
68.07%
13.08.2026
NL2Repo
coding_agent
44.77%
13.08.2026
Cybergym
general_agent
38.70
13.08.2026
DeepSWE
coding_agent
10.04%
13.08.2026
Toolathlon Verified
general_agent
39.92%
26.06.2026
Agents' Last Exam
general_agent
24.53%
13.08.2026
Automation-Bench
general_agent
10.80
13.08.2026
DSBench-FullStack
coding_agent
37.00
13.08.2026
DSBench-Hard
coding_agent
25.80
13.08.2026
MMLU-Pro
knowledge
84.52%
26.06.2026
GPQA Diamond
stem_reasoning
88.28%
26.06.2026
Humanity's Last Exam
stem_reasoning
61.93%
26.06.2026
Apex-Shortlist (no tools)
stem_reasoning
92.66%
26.06.2026
SWE-bench Verified
coding_agent
90.14%
26.06.2026
MMLU
knowledge
95.41%
26.06.2026
MMLU-Redux
knowledge
73.95%
26.06.2026
MMMLU
multilingual
93.45%
26.06.2026
C-Eval
knowledge
90.83%
26.06.2026
SuperGPQA
knowledge
45.72%
26.06.2026
LongBench v2
long_context
44.80%
26.06.2026
MCP-Atlas
general_agent
74.97%
26.06.2026
HMMT Feb 26
stem_reasoning
96.24%
26.06.2026
Apex-Agents
general_agent
75.69%
26.06.2026
Terminal-Bench 2.0
coding_agent
68.31%
26.06.2026
SWE-bench Pro
coding_agent
65.75%
26.06.2026
GDPVal-AA v2
general_agent
76.19%
26.06.2026
CodeForces
stem_reasoning
87.53%
26.06.2026

Model Tree, Spaces and Paper

Model tree for deepseek-ai/DeepSeek-V4-Flash

Adapters

4 models

Finetunes

25 models

Quantizations

127 models

Spaces using deepseek-ai/DeepSeek-V4-Flash 100

Collection including deepseek-ai/DeepSeek-V4-Flash

[

DeepSeek-V4

Collection

10 items • Updated 1 day ago • 900

](https://huggingface.co/collections/deepseek-ai/deepseek-v4)

Paper for deepseek-ai/DeepSeek-V4-Flash

[

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Paper • 2606.19348 • Published Apr 26 • 43

](https://huggingface.co/papers/2606.19348)

Contact

Contact

If you have any questions, please raise an issue or contact us at service@deepseek.com.

Safetensors

Model size

291B params

Tensor type

BF16

·

I64

·

F32

·

F8_E4M3

·

I8

·

Citation

Citation

@misc{deepseekai2026deepseekv4,
      title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
      author={DeepSeek-AI},
      year={2026},
}

License (MIT)

License

This repository and the model weights are licensed under the MIT License.

How to Run Locally

How to Run Locally

Please refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.

For local deployment, we recommend setting the sampling parameters to temperature = 1.0, top_p = 1.0. For the Think Max reasoning mode, we recommend setting the context window to at least 384K tokens.

Chat Template (reasoning modes)

Chat Template

This release does not include a Jinja-format chat template. Instead, we provide a dedicated encoding folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model's text output. Please refer to the encoding folder for full documentation.

A brief example:

from encoding_dsv4 import encode_messages, parse_message_from_completion_text

messages = [
    {"role": "user", "content": "hello"},
    {"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
    {"role": "user", "content": "1+1=?"}
]

## messages -> string
prompt = encode_messages(messages, thinking_mode="thinking")

## string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Pro")
tokens = tokenizer.encode(prompt)

Evaluation Results: Instruct Model

Instruct Model

DeepSeek-V4-Pro and DeepSeek-V4-Flash both support three reasoning effort modes:

Reasoning Mode Characteristics Typical Use Cases Response Format
Non-think Fast, intuitive responses Routine daily tasks, low-risk decisions </think> summary
Think High Conscious logical analysis, slower but more accurate Complex problem-solving, planning <think> thinking </think> summary
Think Max Push reasoning to its fullest extent Exploring the boundary of model reasoning capability Special system prompt + <think> thinking </think> summary

DeepSeek-V4-Pro-Max vs Frontier Models

Benchmark (Metric) Opus-4.6 Max GPT-5.4 xHigh Gemini-3.1-Pro High K2.6 Thinking GLM-5.1 Thinking DS-V4-Pro Max
Knowledge & Reasoning
MMLU-Pro (EM) 89.1 87.5 91.0 87.1 86.0 87.5
SimpleQA-Verified (Pass@1) 46.2 45.3 75.6 36.9 38.1 57.9
Chinese-SimpleQA (Pass@1) 76.4 76.8 85.9 75.9 75.0 84.4
GPQA Diamond (Pass@1) 91.3 93.0 94.3 90.5 86.2 90.1
HLE (Pass@1) 40.0 39.8 44.4 36.4 34.7 37.7
LiveCodeBench (Pass@1) 88.8 - 91.7 89.6 - 93.5
Codeforces (Rating) - 3168 3052 - - 3206
HMMT 2026 Feb (Pass@1) 96.2 97.7 94.7 92.7 89.4 95.2
IMOAnswerBench (Pass@1) 75.3 91.4 81.0 86.0 83.8 89.8
Apex (Pass@1) 34.5 54.1 60.9 24.0 11.5 38.3
Apex Shortlist (Pass@1) 85.9 78.1 89.1 75.5 72.4 90.2
Long Context
MRCR 1M (MMR) 92.9 - 76.3 - - 83.5
CorpusQA 1M (ACC) 71.7 - 53.8 - - 62.0
Agentic
Terminal Bench 2.0 (Acc) 65.4 75.1 68.5 66.7 63.5 67.9
SWE Verified (Resolved) 80.8 - 80.6 80.2 - 80.6
SWE Pro (Resolved) 57.3 57.7 54.2 58.6 58.4 55.4
SWE Multilingual (Resolved) 77.5 - - 76.7 73.3 76.2
BrowseComp (Pass@1) 83.7 82.7 85.9 83.2 79.3 83.4
HLE w/ tools (Pass@1) 53.1 52.0 51.6 54.0 50.4 48.2
GDPval-AA (Elo) 1619 1674 1314 1482 1535 1554
MCPAtlas Public (Pass@1) 73.8 67.2 69.2 66.6 71.8 73.6
Toolathlon (Pass@1) 47.2 54.6 48.8 50.0 40.7 51.8

Comparison across Modes

Benchmark (Metric) V4-Flash Non-Think V4-Flash High V4-Flash Max V4-Pro Non-Think V4-Pro High V4-Pro Max
Knowledge & Reasoning
MMLU-Pro (EM) 83.0 86.4 86.2 82.9 87.1 87.5
SimpleQA-Verified (Pass@1) 23.1 28.9 34.1 45.0 46.2 57.9
Chinese-SimpleQA (Pass@1) 71.5 73.2 78.9 75.8 77.7 84.4
GPQA Diamond (Pass@1) 71.2 87.4 88.1 72.9 89.1 90.1
HLE (Pass@1) 8.1 29.4 34.8 7.7 34.5 37.7
LiveCodeBench (Pass@1) 55.2 88.4 91.6 56.8 89.8 93.5
Codeforces (Rating) - 2816 3052 - 2919 3206
HMMT 2026 Feb (Pass@1) 40.8 91.9 94.8 31.7 94.0 95.2
IMOAnswerBench (Pass@1) 41.9 85.1 88.4 35.3 88.0 89.8
Apex (Pass@1) 1.0 19.1 33.0 0.4 27.4 38.3
Apex Shortlist (Pass@1) 9.3 72.1 85.7 9.2 85.5 90.2
Long Context
MRCR 1M (MMR) 37.5 76.9 78.7 44.7 83.3 83.5
CorpusQA 1M (ACC) 15.5 59.3 60.5 35.6 56.5 62.0
Agentic
Terminal Bench 2.0 (Acc) 49.1 56.6 56.9 59.1 63.3 67.9
SWE Verified (Resolved) 73.7 78.6 79.0 73.6 79.4 80.6
SWE Pro (Resolved) 49.1 52.3 52.6 52.1 54.4 55.4
SWE Multilingual (Resolved) 69.7 70.2 73.3 69.8 74.1 76.2
Br

(table continues in the model card)

Evaluation Results: Base Model

Base Model

Benchmark (Metric) # Shots DeepSeek-V3.2-Base DeepSeek-V4-Flash-Base DeepSeek-V4-Pro-Base
Architecture - MoE MoE MoE
# Activated Params - 37B 13B 49B
# Total Params - 671B 284B 1.6T
World Knowledge
AGIEval (EM) 0-shot 80.1 82.6 83.1
MMLU (EM) 5-shot 87.8 88.7 90.1
MMLU-Redux (EM) 5-shot 87.5 89.4 90.8
MMLU-Pro (EM) 5-shot 65.5 68.3 73.5
MMMLU (EM) 5-shot 87.9 88.8 90.3
C-Eval (EM) 5-shot 90.4 92.1 93.1
CMMLU (EM) 5-shot 88.9 90.4 90.8
MultiLoKo (EM) 5-shot 38.7 42.2 51.1
Simple-QA verified (EM) 25-shot 28.3 30.1 55.2
SuperGPQA (EM) 5-shot 45.0 46.5 53.9
FACTS Parametric (EM) 25-shot 27.1 33.9 62.6
TriviaQA (EM) 5-shot 83.3 82.8 85.6
Language & Reasoning
BBH (EM) 3-shot 87.6 86.9 87.5
DROP (F1) 1-shot 88.2 88.6 88.7
HellaSwag (EM) 0-shot 86.4 85.7 88.0
WinoGrande (EM) 0-shot 78.9 79.5 81.5
CLUEWSC (EM) 5-shot 83.5 82.2 85.2
Code & Math
BigCodeBench (Pass@1) 3-shot 63.9 56.8 59.2
HumanEval (Pass@1) 0-shot 62.8 69.5 76.8
GSM8K (EM) 8-shot 91.1 90.8 92.6
MATH (EM) 4-shot 60.5 57.4 64.5
MGSM (EM) 8-shot 81.3 85.7 84.4
CMath (EM) 3-shot 92.6 93.6 90.9
Long Context
LongBench-V2 (EM) 1-shot 40.2 44.7 51.5

Model Downloads (base/instruct/quantizations)

Model Downloads

Model #Total Params #Activated Params Context Length Precision Download
DeepSeek-V4-Flash-Base 284B 13B 1M FP8 Mixed HuggingFace
DeepSeek-V4-Flash 284B 13B 1M FP4 + FP8 Mixed* HuggingFace
DeepSeek-V4-Pro-Base 1.6T 49B 1M FP8 Mixed HuggingFace
DeepSeek-V4-Pro 1.6T 49B 1M FP4 + FP8 Mixed* HuggingFace

*FP4 + FP8 Mixed: MoE expert parameters use FP4 precision; most other parameters use FP8.

Introduction (DeepSeek-V4 series preview)

Introduction

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

DeepSeek-V4 series incorporate several key upgrades in architecture and optimization:

  1. Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2.
  2. Manifold-Constrained Hyper-Connections (mHC): We incorporate mHC to strengthen conventional residual connections, enhancing stability of signal propagation across layers while preserving model expressivity.
  3. Muon Optimizer: We employ the Muon optimizer for faster convergence and greater training stability.

We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline. The post-training features a two-stage paradigm: independent cultivation of domain-specific experts (through SFT and RL with GRPO), followed by unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model.

DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, significantly advances the knowledge capabilities of open-source models, firmly establishing itself as the best open-source model available today. It achieves top-tier performance in coding benchmarks and significantly bridges the gap with leading closed-source models on reasoning and agentic tasks. Meanwhile, DeepSeek-V4-Flash-Max achieves comparable reasoning performance to the Pro version when given a larger thinking budget, though its smaller parameter scale naturally places it slightly behind on pure knowledge tasks and the most complex agentic workflows.

Architecture

Decoder Block input Embedding vocab 129K · d 4096 Full Attention Sparse Attn 64:1 · dₕ 512 · win 128 ×43 MoE FFN 256 experts · top-6 · +1 shared · dᴻ 2048 MTP Head ×1 speculative layer Final Norm LM Head vocab 129K output
Attention
Sparse Attention (64:1)
MoE
256 experts · top-6 per token
Layers
43
Hidden size
4096
Context
1M tokens
RoPE θ
10K
Parameters
291000M
Active params
13000M

Source: Hugging Face config.json · DeepseekV4ForCausalLM · model repo

Type: MoE
Attention: Hybrid Attention (Compressed Sparse Attention + Heavily Compressed Attention)
Decoder: Autoregressive
MoE: yes (? experts)
Routing: Expert routing
Total parameters 284B
Context length 1M
Extended context 1M
Experts per token 13B activated
Precision FP4 + FP8 Mixed (MoE expert params FP4, most other params FP8)
Vision No
MTP No

Training Pipeline

  1. 1
    pretraining

    Pre-training on 32T+ tokens

    Pre-trained on more than 32T diverse and high-quality tokens. Uses Muon optimizer for faster convergence and greater training stability. MoE architecture with 284B total params (13B activated).

  2. 2
    sft

    Domain-specific expert SFT

    Independent cultivation of domain-specific experts through SFT. Part of first stage of two-stage post-training paradigm.

  3. 3
    rl

    Domain-specific expert RL (GRPO)

    Independent cultivation of domain-specific experts through RL with GRPO. Part of first stage of two-stage post-training paradigm.

  4. 4
    other

    On-policy distillation consolidation

    Unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model. Second stage of two-stage post-training paradigm.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
32T+ token pre-training corpus (diverse, high-quality) pretraining — —
Domain-specific SFT + GRPO RL corpora rl — —

Trend Analysis

24h Change

+0.1%

7d Change

+0.6%

Current

143,962

huggingface

downloads

+1.3%

huggingface

likes

+0.0%

ollama

downloads

+0.8%

huggingface

downloads_all_time

+0.6%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 415,100 pulls daily 01.09.2026
huggingface downloads_all_time 11,080,030 daily 01.09.2026
huggingface followers 143,962 daily 01.09.2026
huggingface likes 2,169 daily 01.09.2026
huggingface downloads 1,727,171 daily 01.09.2026
ollama downloads 412,000 pulls daily 31.08.2026
huggingface downloads_all_time 11,014,089 daily 31.08.2026
huggingface followers 143,846 daily 31.08.2026
huggingface likes 2,169 daily 31.08.2026
huggingface downloads 1,704,533 daily 31.08.2026
ollama downloads 409,200 pulls daily 30.08.2026
huggingface downloads_all_time 10,966,688 daily 30.08.2026
huggingface followers 143,692 daily 30.08.2026
huggingface likes 2,161 daily 30.08.2026
huggingface downloads 1,715,319 daily 30.08.2026
ollama downloads 406,900 pulls daily 29.08.2026
huggingface followers 143,609 daily 29.08.2026
huggingface likes 2,157 daily 29.08.2026
huggingface downloads 1,666,669 daily 29.08.2026
ollama downloads 404,000 pulls daily 28.08.2026
huggingface followers 143,516 daily 28.08.2026
huggingface likes 2,153 daily 28.08.2026
huggingface downloads 1,636,371 daily 28.08.2026
ollama downloads 401,600 pulls daily 27.08.2026
huggingface followers 143,391 daily 27.08.2026
huggingface likes 2,147 daily 27.08.2026
huggingface downloads 1,744,766 daily 27.08.2026
ollama downloads 399,000 pulls daily 26.08.2026
huggingface followers 143,247 daily 26.08.2026
huggingface likes 2,143 daily 26.08.2026
huggingface downloads 1,768,895 daily 26.08.2026
ollama downloads 396,200 pulls daily 25.08.2026
huggingface followers 143,114 daily 25.08.2026
huggingface likes 2,141 daily 25.08.2026
huggingface downloads 1,769,890 daily 25.08.2026
ollama downloads 393,600 pulls daily 24.08.2026
huggingface followers 142,996 daily 24.08.2026
huggingface likes 2,139 daily 24.08.2026
huggingface downloads 1,761,668 daily 24.08.2026

View full metric history →

Related Models