DeepSeek-V4-Flash-0731

DeepSeek

Parameters

304.0B total / 13.0B active

MoE: total / active

Architecture

MoE with DSpark speculative decoding module

Released

31.07.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16, F32, F8_E4M3, I8 DeepSeek en zh

Input Modalities

text

Output Modalities

text

Context (native)

1,000,000 tokens

Context (extended)

1,000,000 tokens

Openness Index Score 100.0/100

About

DeepSeek-V4-Flash-0731 (deepseek-ai/DeepSeek-V4-Flash-0731) is the official release of DeepSeek-V4-Flash, superseding the June preview with substantially enhanced agentic capabilities - released July 31, 2026 under the MIT License with 304B total parameters (13B activated) and a one-million-token context. It has the same model structure as DeepSeek-V4-Flash-DSpark: the V4 hybrid stack (Compressed Sparse Attention (CSA) + Heavily Compressed Attention, Manifold-Constrained Hyper-Connections (mHC), trained with the Muon optimizer) with a DSpark speculative decoding module attached.

Despite its far smaller activated parameter count it outperforms DeepSeek-V4-Pro (Preview) on the card's agentic benchmarks (Terminal Bench 2.1, NL2Repo, Cybergym, DeepSWE, Toolathlon-Verified, Agents' Last Exam, AutomationBench, DSBench) and is broadly competitive with the strongest proprietary models. Post-training follows the V4 two-stage paradigm - domain-specific expert cultivation (SFT + RL with GRPO) then unified consolidation via on-policy distillation. Technical report: arXiv 2606.19348.

Training Data Official release of DeepSeek-V4-Flash, superseding the preview version with substantially enhanced agentic capabilities. Same model structure as DeepSeek-V4-Flash-DSpark with speculative decoding module attached.

Benchmark Scores

Benchmark Score Date
Terminal Bench 2.1
coding_agent
91.24%
31.07.2026
Cybergym
general_agent
76.92%
31.07.2026
DeepSWE
coding_agent
74.83%
31.07.2026
Toolathlon Verified
general_agent
84.83%
31.07.2026
Agents' Last Exam
general_agent
68.87%
31.07.2026
Automation-Bench
general_agent
32.50%
31.07.2026
DSBench-FullStack
coding_agent
78.86%
31.07.2026
DSBench-Hard
coding_agent
73.64%
31.07.2026
Humanity's Last Exam
stem_reasoning
67.61%
13.08.2026
HLE (with tools)
stem_reasoning
72.46%
13.08.2026
DeepSWE 1.1
coding_agent
70.09%
26.08.2026
SWE-bench Pro
coding_agent
70.00%
26.08.2026
NL2Repo-Bench
coding_agent
68.20%
26.08.2026
CoWorkBench
general_agent
45.10
26.08.2026
JobBench
general_agent
42.21%
26.08.2026
IFBench
instruction_following
94.01%
26.08.2026
GPQA Diamond
stem_reasoning
93.38%
26.08.2026
LiveCodeBench v6
stem_reasoning
96.04%
26.08.2026
ApexBench
multimodal_agent
26.20
08.09.2026
NL2Repo
coding_agent
67.54%
31.07.2026

Model Tree & Community Derivatives

Model Tree & Community Derivatives

DeepSeek-V4-Flash-0731 has a substantial community ecosystem of derived models:

Type Count
Adapters 6
Finetunes 31
Merges 1
Quantizations 176
Spaces using this model 31

Collection

Part of the DeepSeek-V4 collection (8 items, updated 11 days ago, 848 followers).

Related Models

  • DeepSeek-V4-Flash-DSpark — Same model structure with speculative decoding module
  • DeepSeek-V4-Pro (Preview) — Pro variant (outperformed by Flash-0731 on benchmarks)

Source: HuggingFace Model Card

BibTeX Citation

Citation

If you use DeepSeek-V4-Flash-0731 in your work, please cite:

@misc{deepseekai2026deepseekv4,
      title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
      author={DeepSeek-AI},
      year={2026},
}

Technical Report

  • arXiv: 2606.19348
  • Title: DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

Source: HuggingFace Model Card

MIT License

License

This repository and the model weights are licensed under the MIT License.

  • Repository: MIT License
  • Commercial use: Allowed (per model DB record commercial_use_allowed = True)
  • Open source: Yes (is_open_source = True)

The MIT License is one of the most permissive open-source licenses, allowing commercial use, modification, distribution, and private use with minimal restrictions.

Source: HuggingFace Model Card

Recommended Sampling Parameters

Recommended Sampling Parameters

General Settings

Parameter Value
temperature 1.0

Agentic vs Non-Agentic Scenarios

Scenario top_p Notes
Agentic tasks 0.95 Use with max reasoning effort
Non-agentic / general 1.0 Standard sampling

Reasoning Effort & Output Length

Reasoning Effort Max Output Length
low Standard
high 384K tokens
max 384K tokens

Evaluation Configuration (Benchmarks)

For Code Agent task evaluation:

  • Agent framework: DeepSeek Harness (minimal mode, to be released)
  • Reasoning effort: max
  • temperature = 1.0, top_p = 0.95

Source: HuggingFace Model Card

Local Deployment Guide

Local Deployment Guide

For local deployment, refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including:

  • Model weight conversion
  • Interactive chat demos

Recommended Sampling Parameters

Parameter Agentic Scenarios Other Scenarios
temperature 1.0 1.0
top_p 0.95 1.0

Output Length

For high and max reasoning effort levels, a maximum output length of 384K tokens is recommended.

Source: HuggingFace Model Card

SGLang Deployment with DSpark

SGLang Deployment with DSpark

Enable DSpark with --speculative-algorithm DSPARK. The target and draft weights come from the same checkpoint, so do not set a separate --speculative-draft-model-path.

Command

sglang serve \
  --trust-remote-code \
  --model-path deepseek-ai/DeepSeek-V4-Flash-0731 \
  --tp 4 \
  --moe-runner-backend flashinfer_mxfp4 \
  --speculative-algorithm DSPARK \
  --mem-fraction-static 0.90 \
  --chunked-prefill-size 4096 \
  --swa-full-tokens-ratio 0.1

Key Parameters

  • --speculative-algorithm DSPARK: Enables DSpark speculative decoding (no separate draft model needed)
  • --tp 4: Tensor parallelism across 4 GPUs
  • --moe-runner-backend flashinfer_mxfp4: MXFP4 quantization backend
  • --mem-fraction-static 0.90: 90% static memory fraction
  • --chunked-prefill-size 4096: Chunked prefill for efficiency
  • --swa-full-tokens-ratio 0.1: Sliding window attention ratio

Cookbook: SGLang cookbook for DeepSeek-V4

Source: HuggingFace Model Card

vLLM Deployment with DSpark Speculative Decoding

vLLM Deployment with DSpark Speculative Decoding

DSpark speculative decoding is enabled with a single flag — add --speculative-config with method: dspark to your vLLM launch command.

Command (single 4×GB300 node)

vllm serve deepseek-ai/DeepSeek-V4-Flash-0731 \
  --trust-remote-code --kv-cache-dtype fp8 --block-size 256 \
  --data-parallel-size 4 --enable-expert-parallel \
  --moe-backend deep_gemm_mega_moe \
  --attention-config '{"use_fp4_indexer_cache": true}' \
  --speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'

Key Parameters

  • --speculative-config: Enables DSpark speculative decoding (7 speculative tokens, greedy draft sampling)
  • --kv-cache-dtype fp8: FP8 KV cache for memory efficiency
  • --data-parallel-size 4: Data parallelism across 4 GPUs
  • --moe-backend deep_gemm_mega_moe: Mega MoE backend for expert computation
  • --attention-config: FP4 indexer cache for attention

Recipe: vLLM recipe for DeepSeek-V4-Flash

Source: HuggingFace Model Card

Reasoning Effort Levels

Reasoning Effort Levels

The reasoning_effort parameter supports three levels that control how much deliberation the model spends before answering:

Level Description
low Minimal deliberation — fast responses
high Moderate deliberation — balanced thinking
max Maximum deliberation — deepest reasoning

Usage

The reasoning effort is set via the reasoning_effort parameter in the encode_messages() function:

prompt = encode_messages(messages, thinking_mode="thinking", reasoning_effort="max")

Output Length Recommendation

For the high and max reasoning effort levels, a maximum output length of 384K tokens is recommended.

Source: HuggingFace Model Card

Chat Template & Encoding

Chat Template & Encoding

This release does not include a Jinja-format chat template. Instead, a dedicated encoding folder with Python scripts and test cases is provided, demonstrating how to:

  1. Encode messages in OpenAI-compatible format into input strings for the model
  2. Parse the model's text output

Please refer to the encoding folder for full documentation.

Example

from encoding_dsv4 import encode_messages, parse_message_from_completion_text

messages = [
    {"role": "user", "content": "hello"},
    {"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
    {"role": "user", "content": "1+1=?"}
]

## messages -> string
prompt = encode_messages(messages, thinking_mode="thinking", reasoning_effort="max")

## string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Flash-0731")
tokens = tokenizer.encode(prompt)

Source: HuggingFace Model Card

Agentic Benchmark Performance

Agentic Benchmark Performance

DeepSeek-V4-Flash-0731 benchmark scores across agentic and coding agent benchmarks:

Benchmark DeepSeek-V4-Flash-0731 DeepSeek-V4-Flash (Preview) DeepSeek-V4-Pro (Preview) GLM-5.2 Opus-4.8
Terminal Bench 2.1 82.7 61.8 72.1 81.0 85.0
NL2Repo 54.2 39.4 38.5 48.9 69.7
Cybergym 76.7 38.7 52.7 - 83.1
DeepSWE 54.4 7.3 12.8 46.2 58.0
Toolathlon-Verified 70.3 49.7 55.9 59.9 76.2
Agents' Last Exam 25.2 15.8 16.5 23.8 25.7
AutomationBench Public 25.1 10.8 12.8 12.9 27.2
DSBench-FullStack † 68.7 37.0 41.8 61.8 71.6
DSBench-Hard † 59.6 25.8 31.1 54.5 71.7

Notes:

  1. For Code Agent tasks, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95.
  2. † DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.

Source: HuggingFace Model Card

Introduction & Key Improvements

Introduction & Key Improvements

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.

DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.

Key highlights:

  • Official release superseding the preview version
  • Substantially enhanced agentic capabilities
  • Same model structure as DeepSeek-V4-Flash-DSpark with speculative decoding module
  • Outperforms DeepSeek-V4-Pro (Preview) despite smaller activated parameter count
  • Broadly competitive with strongest proprietary models

Source: HuggingFace Model Card

Architecture

Decoder Block input Embedding vocab 129K · d 4096 Full Attention Sparse Attn 64:1 · dₕ 512 · win 128 ×43 MoE FFN 256 experts · top-6 · +1 shared · dᴻ 2048 MTP Head ×1 speculative layer Final Norm LM Head vocab 129K output
Attention
Sparse Attention (64:1)
MoE
256 experts · top-6 per token
Layers
43
Hidden size
4096
Context
1M tokens
RoPE θ
10K
Parameters
304000M
Active params
13000M

Source: Hugging Face config.json · DeepseekV4ForCausalLM · model repo

Type: MoE
Attention: Multi-head attention
Decoder: Transformer decoder
MoE: yes (? experts)
Routing: Mixture of Experts with DSpark speculative decoding
Max Output Tokens 393K
Quantization FP8
Inference Frameworks

vLLM, SGLang

Reasoning Effort Levels

low, high, max

Speculative Decoding

DSpark

Speculative Decoding Tokens

7

Training Pipeline

  1. 1
    pretraining

    Pre-training

    32T+ diverse high-quality tokens (V4 family recipe).

  2. 2
    sft

    Domain-specific expert SFT

    Independent cultivation of domain-specific experts through SFT (V4 two-stage post-training paradigm).

  3. 3
    rl

    GRPO RL expert cultivation

    RL with GRPO applied to the domain-specific experts.

  4. 4
    other

    Consolidation + DSpark attachment

    Unified model consolidation via on-policy distillation, integrating distinct proficiencies; DSpark speculative decoding module attached (card Introduction).

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
32T+ token pre-training corpus (diverse, high-quality) pretraining — —
Domain-specific SFT + GRPO RL corpora rl — —

Trend Analysis

24h Change

+0.3%

7d Change

+3.7%

Current

3,851

huggingface

downloads

+1.9%

huggingface

followers

+0.1%

huggingface

downloads_all_time

+3.5%

ollama

downloads

+0.8%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 415,100 pulls daily 01.09.2026
huggingface downloads_all_time 4,874,324 daily 01.09.2026
huggingface followers 143,962 daily 01.09.2026
huggingface likes 3,851 daily 01.09.2026
huggingface downloads 4,650,353 daily 01.09.2026
ollama downloads 412,000 pulls daily 31.08.2026
huggingface downloads_all_time 4,710,705 daily 31.08.2026
huggingface followers 143,846 daily 31.08.2026
huggingface likes 3,841 daily 31.08.2026
huggingface downloads 4,561,861 daily 31.08.2026
ollama downloads 409,200 pulls daily 30.08.2026
huggingface downloads_all_time 4,622,115 daily 30.08.2026
huggingface followers 143,692 daily 30.08.2026
huggingface likes 3,823 daily 30.08.2026
huggingface downloads 4,575,518 daily 30.08.2026
ollama downloads 406,900 pulls daily 29.08.2026
huggingface followers 143,609 daily 29.08.2026
huggingface likes 3,806 daily 29.08.2026
huggingface downloads 4,330,482 daily 29.08.2026
ollama downloads 404,000 pulls daily 28.08.2026
huggingface followers 143,516 daily 28.08.2026
huggingface likes 3,786 daily 28.08.2026
huggingface downloads 3,959,575 daily 28.08.2026
ollama downloads 401,600 pulls daily 27.08.2026
huggingface followers 143,391 daily 27.08.2026
huggingface likes 3,761 daily 27.08.2026
huggingface downloads 3,959,575 daily 27.08.2026
ollama downloads 399,000 pulls daily 26.08.2026
huggingface followers 143,247 daily 26.08.2026
huggingface likes 3,738 daily 26.08.2026
huggingface downloads 3,857,140 daily 26.08.2026
ollama downloads 396,200 pulls daily 25.08.2026
huggingface followers 143,114 daily 25.08.2026
huggingface likes 3,714 daily 25.08.2026
huggingface downloads 3,528,373 daily 25.08.2026
ollama downloads 393,600 pulls daily 24.08.2026
huggingface followers 142,996 daily 24.08.2026
huggingface likes 3,681 daily 24.08.2026
huggingface downloads 3,274,129 daily 24.08.2026
huggingface followers 142,872 daily 23.08.2026
huggingface likes 3,652 daily 23.08.2026
huggingface downloads 3,089,709 daily 23.08.2026
huggingface followers 142,773 daily 22.08.2026
huggingface likes 3,629 daily 22.08.2026
huggingface downloads 2,976,281 daily 22.08.2026
huggingface followers 142,679 daily 21.08.2026
huggingface likes 3,609 daily 21.08.2026
huggingface downloads 2,833,064 daily 21.08.2026
huggingface followers 142,561 daily 20.08.2026
huggingface likes 3,575 daily 20.08.2026

View full metric history →

Related Models