Parameters
304.0B total / 13.0B active
MoE: total / active
Architecture
MoE with DSpark speculative decoding module
Released
31.07.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
1,000,000 tokens
Context (extended)
1,000,000 tokens
About
DeepSeek-V4-Flash-0731 (deepseek-ai/DeepSeek-V4-Flash-0731) is the official release of DeepSeek-V4-Flash, superseding the June preview with substantially enhanced agentic capabilities - released July 31, 2026 under the MIT License with 304B total parameters (13B activated) and a one-million-token context. It has the same model structure as DeepSeek-V4-Flash-DSpark: the V4 hybrid stack (Compressed Sparse Attention (CSA) + Heavily Compressed Attention, Manifold-Constrained Hyper-Connections (mHC), trained with the Muon optimizer) with a DSpark speculative decoding module attached.
Despite its far smaller activated parameter count it outperforms DeepSeek-V4-Pro (Preview) on the card's agentic benchmarks (Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon-Verified, Agents' Last Exam, AutomationBench, DSBench) and is broadly competitive with the strongest proprietary models. Post-training follows the V4 two-stage paradigm - domain-specific expert cultivation (SFT + RL with GRPO) then unified consolidation via on-policy distillation. Technical report: arXiv 2606.19348.
Training Data Official release of DeepSeek-V4-Flash, superseding the preview version with substantially enhanced agentic capabilities. Same model structure as DeepSeek-V4-Flash-DSpark with speculative decoding module attached.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Terminal Bench 2.1
coding_agent
|
91.24%
|
31.07.2026 |
|
Cybergym
general_agent
|
76.92%
|
31.07.2026 |
|
DeepSWE
coding_agent
|
74.83%
|
31.07.2026 |
|
Toolathlon Verified
general_agent
|
84.83%
|
31.07.2026 |
|
Agents' Last Exam
general_agent
|
68.87%
|
31.07.2026 |
|
Automation-Bench
general_agent
|
32.50%
|
31.07.2026 |
|
DSBench-FullStack
coding_agent
|
78.86%
|
31.07.2026 |
|
DSBench-Hard
coding_agent
|
73.64%
|
31.07.2026 |
|
Humanity's Last Exam
stem_reasoning
|
67.61%
|
13.08.2026 |
|
HLE (with tools)
stem_reasoning
|
72.46%
|
13.08.2026 |
|
DeepSWE 1.1
coding_agent
|
70.09%
|
26.08.2026 |
|
SWE-bench Pro
coding_agent
|
70.00%
|
26.08.2026 |
|
NL2Repo-Bench
coding_agent
|
68.20%
|
26.08.2026 |
|
CoWorkBench
general_agent
|
45.10
|
26.08.2026 |
|
JobBench
general_agent
|
42.21%
|
26.08.2026 |
|
IFBench
instruction_following
|
94.01%
|
26.08.2026 |
|
GPQA Diamond
stem_reasoning
|
93.38%
|
26.08.2026 |
|
LiveCodeBench v6
stem_reasoning
|
96.04%
|
26.08.2026 |
|
ApexBench
multimodal_agent
|
26.20
|
08.09.2026 |
|
NL2Repo
coding_agent
|
67.54%
|
31.07.2026 |
Model Tree & Community Derivatives
Model Tree & Community Derivatives
DeepSeek-V4-Flash-0731 has a substantial community ecosystem of derived models:
| Type | Count |
|---|---|
| Adapters | 6 |
| Finetunes | 31 |
| Merges | 1 |
| Quantizations | 176 |
| Spaces using this model | 31 |
Collection
Part of the DeepSeek-V4 collection (8 items, updated 11 days ago, 848 followers).
Related Models
- DeepSeek-V4-Flash-DSpark — Same model structure with speculative decoding module
- DeepSeek-V4-Pro (Preview) — Pro variant (outperformed by Flash-0731 on benchmarks)
Source: HuggingFace Model Card
BibTeX Citation
Citation
If you use DeepSeek-V4-Flash-0731 in your work, please cite:
@misc{deepseekai2026deepseekv4,
title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
author={DeepSeek-AI},
year={2026},
}
Technical Report
- arXiv: 2606.19348
- Title: DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Source: HuggingFace Model Card
MIT License
License
This repository and the model weights are licensed under the MIT License.
- Repository: MIT License
- Commercial use: Allowed (per model DB record
commercial_use_allowed = True) - Open source: Yes (
is_open_source = True)
The MIT License is one of the most permissive open-source licenses, allowing commercial use, modification, distribution, and private use with minimal restrictions.
Source: HuggingFace Model Card
Recommended Sampling Parameters
Recommended Sampling Parameters
General Settings
| Parameter | Value |
|---|---|
temperature |
1.0 |
Agentic vs Non-Agentic Scenarios
| Scenario | top_p |
Notes |
|---|---|---|
| Agentic tasks | 0.95 | Use with max reasoning effort |
| Non-agentic / general | 1.0 | Standard sampling |
Reasoning Effort & Output Length
| Reasoning Effort | Max Output Length |
|---|---|
low |
Standard |
high |
384K tokens |
max |
384K tokens |
Evaluation Configuration (Benchmarks)
For Code Agent task evaluation:
- Agent framework: DeepSeek Harness (minimal mode, to be released)
- Reasoning effort:
max temperature = 1.0, top_p = 0.95
Source: HuggingFace Model Card
Local Deployment Guide
Local Deployment Guide
For local deployment, refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including:
- Model weight conversion
- Interactive chat demos
Recommended Sampling Parameters
| Parameter | Agentic Scenarios | Other Scenarios |
|---|---|---|
temperature |
1.0 | 1.0 |
top_p |
0.95 | 1.0 |
Output Length
For high and max reasoning effort levels, a maximum output length of 384K tokens is recommended.
Source: HuggingFace Model Card
SGLang Deployment with DSpark
SGLang Deployment with DSpark
Enable DSpark with --speculative-algorithm DSPARK. The target and draft weights come from the same checkpoint, so do not set a separate --speculative-draft-model-path.
Command
sglang serve \
--trust-remote-code \
--model-path deepseek-ai/DeepSeek-V4-Flash-0731 \
--tp 4 \
--moe-runner-backend flashinfer_mxfp4 \
--speculative-algorithm DSPARK \
--mem-fraction-static 0.90 \
--chunked-prefill-size 4096 \
--swa-full-tokens-ratio 0.1
Key Parameters
--speculative-algorithm DSPARK: Enables DSpark speculative decoding (no separate draft model needed)--tp 4: Tensor parallelism across 4 GPUs--moe-runner-backend flashinfer_mxfp4: MXFP4 quantization backend--mem-fraction-static 0.90: 90% static memory fraction--chunked-prefill-size 4096: Chunked prefill for efficiency--swa-full-tokens-ratio 0.1: Sliding window attention ratio
Cookbook: SGLang cookbook for DeepSeek-V4
Source: HuggingFace Model Card
vLLM Deployment with DSpark Speculative Decoding
vLLM Deployment with DSpark Speculative Decoding
DSpark speculative decoding is enabled with a single flag — add --speculative-config with method: dspark to your vLLM launch command.
Command (single 4×GB300 node)
vllm serve deepseek-ai/DeepSeek-V4-Flash-0731 \
--trust-remote-code --kv-cache-dtype fp8 --block-size 256 \
--data-parallel-size 4 --enable-expert-parallel \
--moe-backend deep_gemm_mega_moe \
--attention-config '{"use_fp4_indexer_cache": true}' \
--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'
Key Parameters
--speculative-config: Enables DSpark speculative decoding (7 speculative tokens, greedy draft sampling)--kv-cache-dtype fp8: FP8 KV cache for memory efficiency--data-parallel-size 4: Data parallelism across 4 GPUs--moe-backend deep_gemm_mega_moe: Mega MoE backend for expert computation--attention-config: FP4 indexer cache for attention
Recipe: vLLM recipe for DeepSeek-V4-Flash
Source: HuggingFace Model Card
Reasoning Effort Levels
Reasoning Effort Levels
The reasoning_effort parameter supports three levels that control how much deliberation the model spends before answering:
| Level | Description |
|---|---|
low |
Minimal deliberation — fast responses |
high |
Moderate deliberation — balanced thinking |
max |
Maximum deliberation — deepest reasoning |
Usage
The reasoning effort is set via the reasoning_effort parameter in the encode_messages() function:
prompt = encode_messages(messages, thinking_mode="thinking", reasoning_effort="max")
Output Length Recommendation
For the high and max reasoning effort levels, a maximum output length of 384K tokens is recommended.
Source: HuggingFace Model Card
Chat Template & Encoding
Chat Template & Encoding
This release does not include a Jinja-format chat template. Instead, a dedicated encoding folder with Python scripts and test cases is provided, demonstrating how to:
- Encode messages in OpenAI-compatible format into input strings for the model
- Parse the model's text output
Please refer to the encoding folder for full documentation.
Example
from encoding_dsv4 import encode_messages, parse_message_from_completion_text
messages = [
{"role": "user", "content": "hello"},
{"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
{"role": "user", "content": "1+1=?"}
]
## messages -> string
prompt = encode_messages(messages, thinking_mode="thinking", reasoning_effort="max")
## string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Flash-0731")
tokens = tokenizer.encode(prompt)
Source: HuggingFace Model Card
Agentic Benchmark Performance
Agentic Benchmark Performance
DeepSeek-V4-Flash-0731 benchmark scores across agentic and coding agent benchmarks:
| Benchmark | DeepSeek-V4-Flash-0731 | DeepSeek-V4-Flash (Preview) | DeepSeek-V4-Pro (Preview) | GLM-5.2 | Opus-4.8 |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 82.7 | 61.8 | 72.1 | 81.0 | 85.0 |
| NL2Repo | 54.2 | 39.4 | 38.5 | 48.9 | 69.7 |
| Cybergym | 76.7 | 38.7 | 52.7 | - | 83.1 |
| DeepSWE | 54.4 | 7.3 | 12.8 | 46.2 | 58.0 |
| Toolathlon-Verified | 70.3 | 49.7 | 55.9 | 59.9 | 76.2 |
| Agents' Last Exam | 25.2 | 15.8 | 16.5 | 23.8 | 25.7 |
| AutomationBench Public | 25.1 | 10.8 | 12.8 | 12.9 | 27.2 |
| DSBench-FullStack † | 68.7 | 37.0 | 41.8 | 61.8 | 71.6 |
| DSBench-Hard † | 59.6 | 25.8 | 31.1 | 54.5 | 71.7 |
Notes:
- For Code Agent tasks, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the
maxreasoning effort level withtemperature = 1.0, top_p = 0.95. - † DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.
Source: HuggingFace Model Card
Introduction & Key Improvements
Introduction & Key Improvements
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.
DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.
Key highlights:
- Official release superseding the preview version
- Substantially enhanced agentic capabilities
- Same model structure as DeepSeek-V4-Flash-DSpark with speculative decoding module
- Outperforms DeepSeek-V4-Pro (Preview) despite smaller activated parameter count
- Broadly competitive with strongest proprietary models
Source: HuggingFace Model Card
Architecture
- Attention
- Sparse Attention (64:1)
- MoE
- 256 experts · top-6 per token
- Layers
- 43
- Hidden size
- 4096
- Context
- 1M tokens
- RoPE θ
- 10K
- Parameters
- 304000M
- Active params
- 13000M
Source: Hugging Face config.json · DeepseekV4ForCausalLM · model repo
vLLM, SGLang
low, high, max
DSpark
7
Training Pipeline
-
1
pretraining
Pre-training
32T+ diverse high-quality tokens (V4 family recipe).
-
2
sft
Domain-specific expert SFT
Independent cultivation of domain-specific experts through SFT (V4 two-stage post-training paradigm).
-
3
rl
GRPO RL expert cultivation
RL with GRPO applied to the domain-specific experts.
-
4
other
Consolidation + DSpark attachment
Unified model consolidation via on-policy distillation, integrating distinct proficiencies; DSpark speculative decoding module attached (card Introduction).
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| 32T+ token pre-training corpus (diverse, high-quality) | pretraining | — | — | |
| Domain-specific SFT + GRPO RL corpora | rl | — | — |
Linked Resources
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
https://arxiv.org/abs/2606.19348
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
https://huggingface.co/papers/2606.19348
DeepSeek Official Website
https://www.deepseek.com/
DeepSeek Chat
https://chat.deepseek.com/
DeepSeek-V4 Collection
https://huggingface.co/collections/deepseek-ai/deepseek-v4
DeepSeek-V4 Citation
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Trend Analysis
24h Change
+0.3%
7d Change
+3.7%
Current
3,851
downloads
+1.9%
followers
+0.1%
downloads_all_time
+3.5%
downloads
+0.8%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 415,100 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 4,874,324 | daily | 01.09.2026 |
| huggingface | followers | 143,962 | daily | 01.09.2026 |
| huggingface | likes | 3,851 | daily | 01.09.2026 |
| huggingface | downloads | 4,650,353 | daily | 01.09.2026 |
| ollama | downloads | 412,000 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 4,710,705 | daily | 31.08.2026 |
| huggingface | followers | 143,846 | daily | 31.08.2026 |
| huggingface | likes | 3,841 | daily | 31.08.2026 |
| huggingface | downloads | 4,561,861 | daily | 31.08.2026 |
| ollama | downloads | 409,200 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 4,622,115 | daily | 30.08.2026 |
| huggingface | followers | 143,692 | daily | 30.08.2026 |
| huggingface | likes | 3,823 | daily | 30.08.2026 |
| huggingface | downloads | 4,575,518 | daily | 30.08.2026 |
| ollama | downloads | 406,900 pulls | daily | 29.08.2026 |
| huggingface | followers | 143,609 | daily | 29.08.2026 |
| huggingface | likes | 3,806 | daily | 29.08.2026 |
| huggingface | downloads | 4,330,482 | daily | 29.08.2026 |
| ollama | downloads | 404,000 pulls | daily | 28.08.2026 |
| huggingface | followers | 143,516 | daily | 28.08.2026 |
| huggingface | likes | 3,786 | daily | 28.08.2026 |
| huggingface | downloads | 3,959,575 | daily | 28.08.2026 |
| ollama | downloads | 401,600 pulls | daily | 27.08.2026 |
| huggingface | followers | 143,391 | daily | 27.08.2026 |
| huggingface | likes | 3,761 | daily | 27.08.2026 |
| huggingface | downloads | 3,959,575 | daily | 27.08.2026 |
| ollama | downloads | 399,000 pulls | daily | 26.08.2026 |
| huggingface | followers | 143,247 | daily | 26.08.2026 |
| huggingface | likes | 3,738 | daily | 26.08.2026 |
| huggingface | downloads | 3,857,140 | daily | 26.08.2026 |
| ollama | downloads | 396,200 pulls | daily | 25.08.2026 |
| huggingface | followers | 143,114 | daily | 25.08.2026 |
| huggingface | likes | 3,714 | daily | 25.08.2026 |
| huggingface | downloads | 3,528,373 | daily | 25.08.2026 |
| ollama | downloads | 393,600 pulls | daily | 24.08.2026 |
| huggingface | followers | 142,996 | daily | 24.08.2026 |
| huggingface | likes | 3,681 | daily | 24.08.2026 |
| huggingface | downloads | 3,274,129 | daily | 24.08.2026 |
| huggingface | followers | 142,872 | daily | 23.08.2026 |
| huggingface | likes | 3,652 | daily | 23.08.2026 |
| huggingface | downloads | 3,089,709 | daily | 23.08.2026 |
| huggingface | followers | 142,773 | daily | 22.08.2026 |
| huggingface | likes | 3,629 | daily | 22.08.2026 |
| huggingface | downloads | 2,976,281 | daily | 22.08.2026 |
| huggingface | followers | 142,679 | daily | 21.08.2026 |
| huggingface | likes | 3,609 | daily | 21.08.2026 |
| huggingface | downloads | 2,833,064 | daily | 21.08.2026 |
| huggingface | followers | 142,561 | daily | 20.08.2026 |
| huggingface | likes | 3,575 | daily | 20.08.2026 |