DeepSeek-V4-Pro-0813

DeepSeek

Parameters

1.7T total / 49.0B active

MoE: total / active

Architecture

MoE with DSpark speculative decoding module

Released

13.08.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16, I64, F32, F8_E4M3, I8 DeepSeek en zh

Input Modalities

text

Output Modalities

text

Context (native)

1,000,000 tokens

Context (extended)

1,000,000 tokens

Openness Index Score 100.0/100

About

DeepSeek-V4-Pro-0813 (deepseek-ai/DeepSeek-V4-Pro-0813) is the official release of DeepSeek-V4-Pro, superseding the preview version with greatly enhanced agentic capabilities and performance improvements especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure with a DSpark speculative decoding module attached - 1.7T total parameters, 49B activated per token, one-million-token context, released August 13, 2026 under the MIT License.

The V4-Pro stack combines Compressed Sparse Attention (CSA) with Heavily Compressed Attention for long-context efficiency (at 1M context, 27% of single-token FLOPs and 10% of KV cache versus DeepSeek-V3.2) and Manifold-Constrained Hyper-Connections (mHC) for stable deep-stack signal propagation, trained with the Muon optimizer. It outperforms DeepSeek-V4-Pro (Preview) across the card's agentic benchmarks (HLE, Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon-Verified, Agents' Last Exam, AutomationBench, DSBench) and is broadly competitive with the strongest proprietary models. Pre-trained on 32T+ diverse tokens; post-training follows the V4 two-stage paradigm - domain-specific expert cultivation (SFT + RL with GRPO) then consolidation via on-policy distillation. Supports three reasoning effort modes: low, high, max.

Training Data Official release of DeepSeek-V4-Pro, superseding the preview version with greatly enhanced agentic capabilities. Built on DeepSeek-V4-Pro (Preview) model structure with DSpark speculative decoding module attached. Pre-trained on 32T+ diverse tokens. Post-training: two-stage paradigm with independent domain-specific expert cultivation (SFT + RL with GRPO), then unified model consolidation via on-policy distillation. Supports three reasoning effort modes: low, high, max.

Benchmark Scores

Benchmark Score Date
Humanity's Last Exam
stem_reasoning
76.89%
13.08.2026
HLE (with tools)
stem_reasoning
90.47%
13.08.2026
Terminal Bench 2.1
coding_agent
97.01%
13.08.2026
NL2Repo
coding_agent
78.77%
13.08.2026
Cybergym
general_agent
90.28%
13.08.2026
DeepSWE
coding_agent
86.24%
13.08.2026
Toolathlon Verified
general_agent
92.42%
13.08.2026
Agents' Last Exam
general_agent
71.23%
13.08.2026
Automation-Bench
general_agent
47.73%
13.08.2026
DSBench-FullStack
coding_agent
84.83%
13.08.2026
DSBench-Hard
coding_agent
90.20%
13.08.2026
DeepSWE 1.1
coding_agent
82.63%
28.08.2026
GDPVal-AA v2
general_agent
86.84%
28.08.2026

Model Tree, Spaces and Paper

Model tree for deepseek-ai/DeepSeek-V4-Pro-0813

Finetunes

1 model

Quantizations

14 models

Spaces using deepseek-ai/DeepSeek-V4-Pro-0813 14

Collection including deepseek-ai/DeepSeek-V4-Pro-0813

[

DeepSeek-V4

Collection

10 items • Updated 1 day ago • 900

](https://huggingface.co/collections/deepseek-ai/deepseek-v4)

Paper for deepseek-ai/DeepSeek-V4-Pro-0813

[

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Paper • 2606.19348 • Published Apr 26 • 43

](https://huggingface.co/papers/2606.19348)

Contact

Contact

If you have any questions, please raise an issue or contact us at service@deepseek.com.

Safetensors

Model size

1.7T params

Tensor type

BF16

·

I64

·

F32

·

F8_E4M3

·

I8

·

Citation

Citation

@misc{deepseekai2026deepseekv4,
      title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
      author={DeepSeek-AI},
      year={2026},
}

License (MIT)

License

This repository and the model weights are licensed under the MIT License.

How to Run Locally

How to Run Locally

Please refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.

For local deployment, we recommend setting the sampling parameters to temperature = 1.0, with top_p = 0.95 for agentic scenarios and top_p = 1.0 otherwise. For the high and max reasoning effort levels, we recommend a maximum output length of 384K tokens.

How to Run with SGLang

How to Run with SGLang

Enable DSpark with --speculative-algorithm DSPARK and do not set a separate --speculative-draft-model-path as the target and draft weights therefore come from the same checkpoint. See the SGLang cookbook for detailed instructions, benchmarks and other hardwares configurations.

sglang serve \
  --trust-remote-code \
  --model-path deepseek-ai/DeepSeek-V4-Pro-0813 \
  --tp 4 \
  --moe-runner-backend flashinfer_mxfp4 \
  --speculative-algorithm DSPARK \
  --mem-fraction-static 0.90 \
  --chunked-prefill-size 4096 \
  --swa-full-tokens-ratio 0.1 \

How to Run with vLLM

How to Run with vLLM

DSpark speculative decoding is enabled with a single flag — add --speculative-config with method: dspark to your vLLM launch command:

--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'

For example, the command below serves the model with vLLM on a single 4×GB300 node. See the vLLM recipe for detailed instructions and other hardware configurations.

vllm serve deepseek-ai/DeepSeek-V4-Pro-0813 \
  --trust-remote-code --kv-cache-dtype fp8 --block-size 256 \
  --data-parallel-size 4 --enable-expert-parallel \
  --moe-backend deep_gemm_mega_moe \
  --attention-config '{"use_fp4_indexer_cache": true}' \
  --speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'

Chat Template (reasoning effort modes)

Chat Template

This release does not include a Jinja-format chat template. Instead, we provide a dedicated encoding folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model's text output. Please refer to the encoding folder for full documentation.

The reasoning_effort parameter now supports three levels — low, high, and max — which control how much deliberation the model spends before answering.

A brief example:

from encoding_dsv4 import encode_messages, parse_message_from_completion_text

messages = [
    {"role": "user", "content": "hello"},
    {"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
    {"role": "user", "content": "1+1=?"}
]

## messages -> string
prompt = encode_messages(messages, thinking_mode="thinking", reasoning_effort="max")

## string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Pro-0813")
tokens = tokenizer.encode(prompt)

Introduction (official release, benchmark table)

Introduction

DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.

DeepSeek-V4-Pro-0813 outperforms DeepSeek-V4-Pro (Preview) on the benchmarks listed below, and is broadly competitive with the strongest proprietary models available.

Benchmark DeepSeek-V4-Pro-0813 DeepSeek-V4-Flash-0731 DeepSeek-V4-Pro (Preview) DeepSeek-V4-Flash (Preview) GLM-5.2 Kimi K3 Opus-4.8 Fable-5 (w/ fallback)
HLE (wo / w tools) 42.7 / 60.0 37.8 / 51.5 37.7 / 48.2 34.8 / 45.1 40.5 / 54.7 43.5 / 56.0 49.8 / 57.9 53.3 / 63.0
Terminal Bench 2.1 87.9 82.7 72.1 61.8 81.0 88.3 85.0 88.0
NL2Repo 61.5 54.2 38.5 39.4 48.9 - 69.7 -
Cybergym 83.3 76.7 52.7 38.7 - 80.0 78.3 83.1
DeepSWE 62.7 54.4 12.8 7.3 46.2 67.5 58.0 70.0
Toolathlon-Verified 74.1 70.3 55.9 49.7 59.9 76.5 76.2 77.9
Agents' Last Exam 25.7 25.2 16.5 15.8 23.8 27.6 25.7 -
AutomationBench (Public) 31.8 25.1 12.8 10.8 12.9 30.8 27.2 29.1
DSBench-FullStack † 71.1 68.7 41.8 37.0 61.8 73.7 71.6 77.2
DSBench-Hard † 67.2 59.6 31.1 25.8 54.5 63.0 71.7 68.3

Notes:

  1. For the code-agent tasks among the public benchmarks above, DeepSeek-V4-Pro-0813 is evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95.
  2. † DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.

Architecture

Decoder Block input Embedding vocab 129K · d 7168 Full Attention Sparse Attn 128:1 · dₕ 512 · win 128 ×61 MoE FFN 384 experts · top-6 · +1 shared · dᴻ 3072 MTP Head ×1 speculative layer Final Norm LM Head vocab 129K output
Attention
Sparse Attention (128:1)
MoE
384 experts · top-6 per token
Layers
61
Hidden size
7168
Context
1M tokens
RoPE θ
10K
Parameters
1700000M
Active params
49000M

Source: Hugging Face config.json · DeepseekV4ForCausalLM · model repo

Type: MoE
Attention: Hybrid Attention (Compressed Sparse Attention + Heavily Compressed Attention)
Decoder: Autoregressive
MoE: yes (? experts)
Routing: Expert routing
Total parameters 1.7T (with DSpark speculative decoding module)
Context length 1M
Extended context 1M
Experts per token 49B activated
Precision FP4 + FP8 Mixed (MoE expert params FP4, most other params FP8)
Vision No
MTP No
Speculative Decoding

DSpark (num_speculative_tokens=7)

Training Pipeline

  1. 1
    pretraining

    Pre-training on 32T+ tokens

    Pre-trained on more than 32T diverse and high-quality tokens. Uses Muon optimizer for faster convergence and greater training stability. MoE architecture with 1.7T total params (49B activated). DSpark speculative decoding module attached.

  2. 2
    sft

    Domain-specific expert SFT

    Independent cultivation of domain-specific experts through SFT. Part of first stage of two-stage post-training paradigm.

  3. 3
    rl

    Domain-specific expert RL (GRPO)

    Independent cultivation of domain-specific experts through RL with GRPO. Part of first stage of two-stage post-training paradigm.

  4. 4
    other

    On-policy distillation consolidation

    Unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model. Second stage of two-stage post-training paradigm.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
32T+ token pre-training corpus (diverse, high-quality) pretraining — —
Domain-specific SFT + GRPO RL corpora rl — —

Trend Analysis

24h Change

+3.1%

7d Change

+85.8%

Current

138,833

huggingface

likes

+0.1%

huggingface

followers

+0.1%

huggingface

downloads_all_time

+3.1%

ollama

downloads

+0.8%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 372,700 pulls daily 01.09.2026
huggingface downloads_all_time 138,833 daily 01.09.2026
huggingface followers 143,962 daily 01.09.2026
huggingface likes 793 daily 01.09.2026
huggingface downloads 138,833 daily 01.09.2026
ollama downloads 369,900 pulls daily 31.08.2026
huggingface downloads_all_time 134,723 daily 31.08.2026
huggingface followers 143,846 daily 31.08.2026
huggingface likes 792 daily 31.08.2026
huggingface downloads 134,723 daily 31.08.2026
ollama downloads 367,500 pulls daily 30.08.2026
huggingface downloads_all_time 127,009 daily 30.08.2026
huggingface followers 143,692 daily 30.08.2026
huggingface likes 786 daily 30.08.2026
huggingface downloads 127,009 daily 30.08.2026
ollama downloads 365,600 pulls daily 29.08.2026
huggingface followers 143,609 daily 29.08.2026
huggingface likes 782 daily 29.08.2026
huggingface downloads 111,121 daily 29.08.2026
ollama downloads 363,300 pulls daily 28.08.2026
huggingface followers 143,516 daily 28.08.2026
huggingface likes 776 daily 28.08.2026
huggingface downloads 90,822 daily 28.08.2026
ollama downloads 361,300 pulls daily 27.08.2026
huggingface followers 143,391 daily 27.08.2026
huggingface likes 768 daily 27.08.2026
huggingface downloads 90,822 daily 27.08.2026
ollama downloads 358,900 pulls daily 26.08.2026
huggingface followers 143,247 daily 26.08.2026
huggingface likes 762 daily 26.08.2026
huggingface downloads 85,230 daily 26.08.2026
ollama downloads 356,700 pulls daily 25.08.2026
huggingface followers 143,114 daily 25.08.2026
huggingface likes 758 daily 25.08.2026
huggingface downloads 74,707 daily 25.08.2026
ollama downloads 354,700 pulls daily 24.08.2026
huggingface followers 142,996 daily 24.08.2026
huggingface likes 748 daily 24.08.2026
huggingface downloads 63,058 daily 24.08.2026

View full metric history →

Related Models