Parameters
1.7T total / 49.0B active
MoE: total / active
Architecture
MoE with DSpark speculative decoding module
Released
13.08.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
1,000,000 tokens
Context (extended)
1,000,000 tokens
About
DeepSeek-V4-Pro-0813 (deepseek-ai/DeepSeek-V4-Pro-0813) is the official release of DeepSeek-V4-Pro, superseding the preview version with greatly enhanced agentic capabilities and performance improvements especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure with a DSpark speculative decoding module attached - 1.7T total parameters, 49B activated per token, one-million-token context, released August 13, 2026 under the MIT License.
The V4-Pro stack combines Compressed Sparse Attention (CSA) with Heavily Compressed Attention for long-context efficiency (at 1M context, 27% of single-token FLOPs and 10% of KV cache versus DeepSeek-V3.2) and Manifold-Constrained Hyper-Connections (mHC) for stable deep-stack signal propagation, trained with the Muon optimizer. It outperforms DeepSeek-V4-Pro (Preview) across the card's agentic benchmarks (HLE, Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon-Verified, Agents' Last Exam, AutomationBench, DSBench) and is broadly competitive with the strongest proprietary models. Pre-trained on 32T+ diverse tokens; post-training follows the V4 two-stage paradigm - domain-specific expert cultivation (SFT + RL with GRPO) then consolidation via on-policy distillation. Supports three reasoning effort modes: low, high, max.
Training Data Official release of DeepSeek-V4-Pro, superseding the preview version with greatly enhanced agentic capabilities. Built on DeepSeek-V4-Pro (Preview) model structure with DSpark speculative decoding module attached. Pre-trained on 32T+ diverse tokens. Post-training: two-stage paradigm with independent domain-specific expert cultivation (SFT + RL with GRPO), then unified model consolidation via on-policy distillation. Supports three reasoning effort modes: low, high, max.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Humanity's Last Exam
stem_reasoning
|
76.89%
|
13.08.2026 |
|
HLE (with tools)
stem_reasoning
|
90.47%
|
13.08.2026 |
|
Terminal Bench 2.1
coding_agent
|
97.01%
|
13.08.2026 |
|
NL2Repo
coding_agent
|
78.77%
|
13.08.2026 |
|
Cybergym
general_agent
|
90.28%
|
13.08.2026 |
|
DeepSWE
coding_agent
|
86.24%
|
13.08.2026 |
|
Toolathlon Verified
general_agent
|
92.42%
|
13.08.2026 |
|
Agents' Last Exam
general_agent
|
71.23%
|
13.08.2026 |
|
Automation-Bench
general_agent
|
47.73%
|
13.08.2026 |
|
DSBench-FullStack
coding_agent
|
84.83%
|
13.08.2026 |
|
DSBench-Hard
coding_agent
|
90.20%
|
13.08.2026 |
|
DeepSWE 1.1
coding_agent
|
82.63%
|
28.08.2026 |
|
GDPVal-AA v2
general_agent
|
86.84%
|
28.08.2026 |
Model Tree, Spaces and Paper
Model tree for deepseek-ai/DeepSeek-V4-Pro-0813
Finetunes
Quantizations
Spaces using deepseek-ai/DeepSeek-V4-Pro-0813 14
Collection including deepseek-ai/DeepSeek-V4-Pro-0813
[
DeepSeek-V4
Collection
10 items • Updated 1 day ago • 900
](https://huggingface.co/collections/deepseek-ai/deepseek-v4)
Paper for deepseek-ai/DeepSeek-V4-Pro-0813
[
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Paper • 2606.19348 • Published Apr 26 • 43
](https://huggingface.co/papers/2606.19348)
Contact
Contact
If you have any questions, please raise an issue or contact us at service@deepseek.com.
Model size
1.7T params
Tensor type
BF16
·
I64
·
F32
·
F8_E4M3
·
I8
·
Citation
Citation
@misc{deepseekai2026deepseekv4,
title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
author={DeepSeek-AI},
year={2026},
}
License (MIT)
License
This repository and the model weights are licensed under the MIT License.
How to Run Locally
How to Run Locally
Please refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.
For local deployment, we recommend setting the sampling parameters to temperature = 1.0, with top_p = 0.95 for agentic scenarios and top_p = 1.0 otherwise. For the high and max reasoning effort levels, we recommend a maximum output length of 384K tokens.
How to Run with SGLang
How to Run with SGLang
Enable DSpark with --speculative-algorithm DSPARK and do not set a separate --speculative-draft-model-path as the target and draft weights therefore come from the same checkpoint. See the SGLang cookbook for detailed instructions, benchmarks and other hardwares configurations.
sglang serve \
--trust-remote-code \
--model-path deepseek-ai/DeepSeek-V4-Pro-0813 \
--tp 4 \
--moe-runner-backend flashinfer_mxfp4 \
--speculative-algorithm DSPARK \
--mem-fraction-static 0.90 \
--chunked-prefill-size 4096 \
--swa-full-tokens-ratio 0.1 \
How to Run with vLLM
How to Run with vLLM
DSpark speculative decoding is enabled with a single flag — add --speculative-config with method: dspark to your vLLM launch command:
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'
For example, the command below serves the model with vLLM on a single 4×GB300 node. See the vLLM recipe for detailed instructions and other hardware configurations.
vllm serve deepseek-ai/DeepSeek-V4-Pro-0813 \
--trust-remote-code --kv-cache-dtype fp8 --block-size 256 \
--data-parallel-size 4 --enable-expert-parallel \
--moe-backend deep_gemm_mega_moe \
--attention-config '{"use_fp4_indexer_cache": true}' \
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'
Chat Template (reasoning effort modes)
Chat Template
This release does not include a Jinja-format chat template. Instead, we provide a dedicated encoding folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model's text output. Please refer to the encoding folder for full documentation.
The reasoning_effort parameter now supports three levels — low, high, and max — which control how much deliberation the model spends before answering.
A brief example:
from encoding_dsv4 import encode_messages, parse_message_from_completion_text
messages = [
{"role": "user", "content": "hello"},
{"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
{"role": "user", "content": "1+1=?"}
]
## messages -> string
prompt = encode_messages(messages, thinking_mode="thinking", reasoning_effort="max")
## string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Pro-0813")
tokens = tokenizer.encode(prompt)
Introduction (official release, benchmark table)
Introduction
DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.
DeepSeek-V4-Pro-0813 outperforms DeepSeek-V4-Pro (Preview) on the benchmarks listed below, and is broadly competitive with the strongest proprietary models available.
| Benchmark | DeepSeek-V4-Pro-0813 | DeepSeek-V4-Flash-0731 | DeepSeek-V4-Pro (Preview) | DeepSeek-V4-Flash (Preview) | GLM-5.2 | Kimi K3 | Opus-4.8 | Fable-5 (w/ fallback) |
|---|---|---|---|---|---|---|---|---|
| HLE (wo / w tools) | 42.7 / 60.0 | 37.8 / 51.5 | 37.7 / 48.2 | 34.8 / 45.1 | 40.5 / 54.7 | 43.5 / 56.0 | 49.8 / 57.9 | 53.3 / 63.0 |
| Terminal Bench 2.1 | 87.9 | 82.7 | 72.1 | 61.8 | 81.0 | 88.3 | 85.0 | 88.0 |
| NL2Repo | 61.5 | 54.2 | 38.5 | 39.4 | 48.9 | - | 69.7 | - |
| Cybergym | 83.3 | 76.7 | 52.7 | 38.7 | - | 80.0 | 78.3 | 83.1 |
| DeepSWE | 62.7 | 54.4 | 12.8 | 7.3 | 46.2 | 67.5 | 58.0 | 70.0 |
| Toolathlon-Verified | 74.1 | 70.3 | 55.9 | 49.7 | 59.9 | 76.5 | 76.2 | 77.9 |
| Agents' Last Exam | 25.7 | 25.2 | 16.5 | 15.8 | 23.8 | 27.6 | 25.7 | - |
| AutomationBench (Public) | 31.8 | 25.1 | 12.8 | 10.8 | 12.9 | 30.8 | 27.2 | 29.1 |
| DSBench-FullStack † | 71.1 | 68.7 | 41.8 | 37.0 | 61.8 | 73.7 | 71.6 | 77.2 |
| DSBench-Hard † | 67.2 | 59.6 | 31.1 | 25.8 | 54.5 | 63.0 | 71.7 | 68.3 |
Notes:
- For the code-agent tasks among the public benchmarks above, DeepSeek-V4-Pro-0813 is evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the
maxreasoning effort level withtemperature = 1.0, top_p = 0.95. - † DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.
Architecture
- Attention
- Sparse Attention (128:1)
- MoE
- 384 experts · top-6 per token
- Layers
- 61
- Hidden size
- 7168
- Context
- 1M tokens
- RoPE θ
- 10K
- Parameters
- 1700000M
- Active params
- 49000M
Source: Hugging Face config.json · DeepseekV4ForCausalLM · model repo
DSpark (num_speculative_tokens=7)
Training Pipeline
-
1
pretraining
Pre-training on 32T+ tokens
Pre-trained on more than 32T diverse and high-quality tokens. Uses Muon optimizer for faster convergence and greater training stability. MoE architecture with 1.7T total params (49B activated). DSpark speculative decoding module attached.
-
2
sft
Domain-specific expert SFT
Independent cultivation of domain-specific experts through SFT. Part of first stage of two-stage post-training paradigm.
-
3
rl
Domain-specific expert RL (GRPO)
Independent cultivation of domain-specific experts through RL with GRPO. Part of first stage of two-stage post-training paradigm.
-
4
other
On-policy distillation consolidation
Unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model. Second stage of two-stage post-training paradigm.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| 32T+ token pre-training corpus (diverse, high-quality) | pretraining | — | — | |
| Domain-specific SFT + GRPO RL corpora | rl | — | — |
Linked Resources
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
https://arxiv.org/abs/2606.19348
DeepSeek Homepage
https://www.deepseek.com/
DeepSeek Chat
https://chat.deepseek.com/
vLLM Recipe
https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Pro
SGLang Cookbook
https://docs.sglang.io/cookbook/autoregressive/DeepSeek/DeepSeek-V4#hw=gb300&variant=flash-official&quant=fp4&strategy=low-latency&nodes=single
DeepSeek-V4 Citation
https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
DeepSeek-V4 Collection
https://huggingface.co/collections/deepseek-ai/deepseek-v4
Trend Analysis
24h Change
+3.1%
7d Change
+85.8%
Current
138,833
likes
+0.1%
followers
+0.1%
downloads_all_time
+3.1%
downloads
+0.8%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 372,700 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 138,833 | daily | 01.09.2026 |
| huggingface | followers | 143,962 | daily | 01.09.2026 |
| huggingface | likes | 793 | daily | 01.09.2026 |
| huggingface | downloads | 138,833 | daily | 01.09.2026 |
| ollama | downloads | 369,900 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 134,723 | daily | 31.08.2026 |
| huggingface | followers | 143,846 | daily | 31.08.2026 |
| huggingface | likes | 792 | daily | 31.08.2026 |
| huggingface | downloads | 134,723 | daily | 31.08.2026 |
| ollama | downloads | 367,500 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 127,009 | daily | 30.08.2026 |
| huggingface | followers | 143,692 | daily | 30.08.2026 |
| huggingface | likes | 786 | daily | 30.08.2026 |
| huggingface | downloads | 127,009 | daily | 30.08.2026 |
| ollama | downloads | 365,600 pulls | daily | 29.08.2026 |
| huggingface | followers | 143,609 | daily | 29.08.2026 |
| huggingface | likes | 782 | daily | 29.08.2026 |
| huggingface | downloads | 111,121 | daily | 29.08.2026 |
| ollama | downloads | 363,300 pulls | daily | 28.08.2026 |
| huggingface | followers | 143,516 | daily | 28.08.2026 |
| huggingface | likes | 776 | daily | 28.08.2026 |
| huggingface | downloads | 90,822 | daily | 28.08.2026 |
| ollama | downloads | 361,300 pulls | daily | 27.08.2026 |
| huggingface | followers | 143,391 | daily | 27.08.2026 |
| huggingface | likes | 768 | daily | 27.08.2026 |
| huggingface | downloads | 90,822 | daily | 27.08.2026 |
| ollama | downloads | 358,900 pulls | daily | 26.08.2026 |
| huggingface | followers | 143,247 | daily | 26.08.2026 |
| huggingface | likes | 762 | daily | 26.08.2026 |
| huggingface | downloads | 85,230 | daily | 26.08.2026 |
| ollama | downloads | 356,700 pulls | daily | 25.08.2026 |
| huggingface | followers | 143,114 | daily | 25.08.2026 |
| huggingface | likes | 758 | daily | 25.08.2026 |
| huggingface | downloads | 74,707 | daily | 25.08.2026 |
| ollama | downloads | 354,700 pulls | daily | 24.08.2026 |
| huggingface | followers | 142,996 | daily | 24.08.2026 |
| huggingface | likes | 748 | daily | 24.08.2026 |
| huggingface | downloads | 63,058 | daily | 24.08.2026 |