Parameters
763.0B total / 16.0B active
MoE: total / active
Architecture
MoE
Released
10.09.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
1,000,000 tokens
Context (extended)
1,000,000 tokens
About
DeepSeek-V4.1-Flash is DeepSeek's MIT-licensed multimodal Mixture-of-Experts model with 552B backbone parameters (763B total) and a 1M-token context window. Maintained by DeepSeek, it natively processes images and text (DeepSeek-ViT vision encoder plus a two-layer MLP projector) and generates text autoregressively, with English and Chinese as primary languages.
Its defining novelty is extreme KV-cache compression for agentic workloads: the global KV cache shrinks to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash and 1/437 of DeepSeek-V1 — while activating only 8B parameters per token during prefill and 16B during decode.
- Causal Encoder-Decoder (CED): a 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder; the decoder's global KV cache is projected from the final encoder hidden states rather than from each decoder layer's own states.
- Compressed Sparse Attention 2 (CSA2): every attention layer runs in one of three static modes (Full, Reindex, Reuse), sharing main KV and indexer K across layers and reusing Top-K sparse-attention indices; a Hierarchical Sparse Indexer bounds deeper indexer cost independently of context length.
- FP4 main KV caching (E2M1 format, one E4M3 scale per 16 channels) and SWA Bounded Replay, which reconstructs missing sliding-window-attention KV states by replaying only the most recent n_win tokens — cutting the persistent KV footprint to about 1/8 of DeepSeek-V4-Flash without persisting SWA KV to SSD.
- Sparse MoE: 1 shared expert plus 384 routed experts per layer with 6 routed experts activated per token; Engram conditional memory (196B parameters, token-based lookup), DSpark speculative decoding (semi-autoregressive drafts with confidence-scheduled verification) and Single-Pass mHC with an efficient Mega-mHC kernel.
Pre-training used a 45T-token multimodal corpus with sparse attention trained at 64K sequence length, extending context to 1M tokens at 34T tokens. Post-training follows the SFT → RL → on-policy distillation (OPD) paradigm, with large-scale automated synthesis of agent tasks and environments; a continuously controllable reasoning_effort setting (integer 1–100) trades inference cost for accuracy.
At maximum reasoning effort the model sets records on DeepSWE v1.1 (74.2) and Terminal-Bench 2.1 (90.6) and leads on CyberGym (88.1), AutomationBench (54.8) and Agent's Last Exam (31.8). Recommended sampling: temperature 1.0, top_p 0.95, context window 1M tokens, max_tokens ≥ 256K.
Training Data Trained from scratch on a 45T-token multimodal corpus; sparse attention trained at 64K sequence length, context extended to 1M tokens at 34T tokens. Post-training: SFT -> RL -> on-policy distillation (OPD) with large-scale automated synthesis of agent tasks and environments.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
GPQA Diamond
stem_reasoning
|
93.57%
|
10.09.2026 |
|
Humanity's Last Exam
stem_reasoning
|
65.72%
|
10.09.2026 |
|
MathArena Apex 2025
stem_reasoning
|
51.04%
|
10.09.2026 |
|
Terminal Bench 2.1
coding_agent
|
100.00%
|
10.09.2026 |
|
Terminal-Bench 3.0
coding_agent
|
84.67%
|
10.09.2026 |
|
Terminal-Bench 4.0
coding_agent
|
100.00%
|
10.09.2026 |
|
DeepSWE 1.1
coding_agent
|
100.00%
|
10.09.2026 |
|
ProgramBench
coding_agent
|
25.11%
|
10.09.2026 |
|
NL2Repo-Bench
coding_agent
|
88.70%
|
10.09.2026 |
|
Cybergym
general_agent
|
100.00%
|
10.09.2026 |
|
SEC-Bench Pro
general_agent
|
100.00%
|
10.09.2026 |
|
HLE (with tools)
stem_reasoning
|
98.73%
|
10.09.2026 |
|
Automation-Bench
general_agent
|
100.00%
|
10.09.2026 |
|
Agents' Last Exam
general_agent
|
100.00%
|
10.09.2026 |
|
Chartography
multimodal_agent
|
100.00%
|
10.09.2026 |
|
BabyVision
vision_language
|
100.00%
|
10.09.2026 |
|
ZEROBench
vision_language
|
100.00%
|
10.09.2026 |
|
CodeForces
stem_reasoning
|
100.00%
|
10.09.2026 |
Model tree, quantizations and ecosystem
- Finetunes: 3 models
- Quantizations: 23 models (llama.cpp, LM Studio, Jan, Ollama targets)
- Spaces: 4 spaces using this model
- Collection: part of the DeepSeek-V4 collection (10 items)
- Base model for community GGUF variants (e.g. FP8/8-bit/NVFP4 builds)
Citation
@misc{deepseekai2026deepseekv41flash,
title={DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression},
author={DeepSeek-AI},
year={2026},
}
License: MIT
The repository and the model weights are licensed under the MIT License (see the LICENSE file in the repository).
Agentic evaluation setup and scaffold comparison
Code agent benchmarks (Terminal-Bench 2.1/3.0/4.0, DeepSWE v1.1, NL2Repo-Bench, ProgramBench) are evaluated with the Minimal mode of DeepSeek Harness and a 1M-token context window; the mini-SWE harness is used for DeepSWE v1.1 and the Claude Code harness for SEC-Bench Pro. Visual agent benchmarks use the Claude Code harness with 512K context. Scaffold comparison (max effort):
| Benchmark | Claude Code | Codex | OpenCode | Pi | mini-SWE | DSH Minimal | DSH Standard | DSH PTC |
|---|---|---|---|---|---|---|---|---|
| DeepSWE v1.1 | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 |
| Terminal-Bench 2.1 | 88.0 | 84.1 | 85.0 | 86.1 | 90.3 | 90.6 | 85.8 | 85.8 |
All scaffolds: N=8 samples (DeepSWE) / N=3 (Terminal-Bench 2.1), temperature=1.0, top_p=0.95, 1M-token context limit, max_steps=500, no network access for Terminal-Bench 2.1. The evaluation/ folder reproduces the DeepSWE v1.1 results.
Prompt encoding and local inference
This release does not include a Jinja-format chat template. The encoding/ folder contains a self-contained Python reference implementation (encoding.py) with test cases for multi-turn conversations, tool calling, thinking mode, numeric reasoning effort, mid-conversation system messages, and interleaved image content.
For production use, deepseek-recipe (github.com/deepseek-ai/deepseek-recipe) provides the same prompt format as maintained Rust libraries with Python bindings, covering Messages, Chat Completions, and Responses API requests. The inference/ folder documents weight conversion and local inference.
Recommended sampling parameters
| Parameter | Value |
|---|---|
temperature |
1.0 |
top_p |
0.95 or 1.0 |
context_window |
1M tokens |
max_tokens |
>= 256K |
Evaluations use temperature=1.0, top_p=0.95.
Continuously controllable reasoning effort
The model supports a continuously controllable reasoning effort setting - an integer from 1 to 100 - that trades inference cost for accuracy. All published instruct results use the maximum effort setting (reasoning_effort=100).
Post-training: SFT -> RL -> on-policy distillation
The post-training recipe follows the standard SFT -> RL -> on-policy distillation (OPD) paradigm without algorithmic modifications. All substantive changes lie in the data pipeline: large-scale automated synthesis of agent tasks and environments with progressive scaling of data, tasks, and rollouts.
Pre-training corpus and context extension
DeepSeek-V4.1-Flash is trained from scratch on a multimodal corpus comprising 45T tokens. Sparse attention is trained at a sequence length of 64K, and context is extended to 1M tokens at the 34T-token mark.
Vision encoder and multimodal pipeline
A vision encoder (DeepSeek-ViT, trained from scratch with 2D-RoPE and 3x3 pixel-unshuffle downsampling) and a two-layer MLP projector convert images into visual embeddings, which are processed jointly with text embeddings from the start of language-model pre-training. Base-model multimodal results: MMMU-Pro 56.5, CVBench 77.9, DocVQA 95.6, RefCOCO-avg 86.0.
MoE design, memory and speculative decoding
- Sparse MoE: 1 shared expert + 384 routed experts per MoE layer, activating 6 routed experts per token
- Engram conditional memory: 196B parameters, sparsely accessed via token-based lookup
- DSPark speculative decoding: semi-autoregressive draft generation with confidence-scheduled verification
- Single-Pass mHC: revised residual-stream mixing with an efficient Mega-mHC kernel
CSA2 attention and KV-cache compression
Compressed Sparse Attention 2 (CSA2) assigns each attention layer one of three static modes - Full, Reindex, or Reuse - to share main KV and indexer K across layers and reuse Top-K sparse-attention indices. In the decoder, a Hierarchical Sparse Indexer restricts later indexing layers to a candidate pool constructed by the first Full Mode layer, bounding deeper indexer cost independently of context length.
Combined with FP4 main KV caching (E2M1 format, one E4M3 scale per 16 channels), the global KV cache footprint drops to 890 bytes per token - roughly 1/4 of DeepSeek-V4-Flash. SWA Bounded Replay reconstructs missing sliding-window-attention KV states by replaying only the most recent n_win tokens, avoiding the need to persist SWA KV to SSD and reducing the persistent KV cache footprint to about 1/8 of DeepSeek-V4-Flash.
Causal Encoder-Decoder (CED) layout
DeepSeek-V4.1-Flash organizes its 40 Transformer layers as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially improving cost efficiency for input-heavy agentic workloads.
DeepSeek-V4.1-Flash at a glance
- Multimodal MoE with 552B backbone parameters (763B total), MIT-licensed
- 1M-token context window, native image + text input, autoregressive text output
- 8B parameters active per token during prefill, 16B during decode
- Global KV cache: 890 bytes per token (~1/4 of DeepSeek-V4-Flash, ~1/437 of DeepSeek-V1)
- Records on DeepSWE v1.1 (74.2) and Terminal-Bench 2.1 (90.6) at max reasoning effort
- Continuously controllable reasoning effort (integer 1-100)
Architecture
- Attention
- Sparse Attention (64:1)
- MoE
- 384 experts · top-6 per token
- Layers
- 40
- Hidden size
- 5120
- Context
- 1M tokens
- RoPE θ
- 10K
- Parameters
- 763000M
- Active params
- 16000M
Source: Hugging Face config.json · DeepseekV41ForCausalLM · model repo
DSPark (semi-autoregressive draft generation with confidence-scheduled verification)
Training Pipeline
-
1
pretraining
Multimodal pre-training from scratch
45T-token multimodal corpus (text + images); sparse attention trained at a sequence length of 64K.
-
2
cpt
Context extension to 1M tokens
Context extended to 1M tokens at the 34T-token mark of pre-training.
-
3
sft
Supervised fine-tuning
Standard SFT stage without algorithmic modifications; substantive changes lie in the data pipeline (large-scale automated synthesis of agent tasks and environments).
-
4
rl
Reinforcement learning
Standard RL stage of the SFT -> RL -> OPD post-training recipe.
-
5
other
On-policy distillation (OPD)
Final consolidation stage; on-policy distillation with progressive scaling of data, tasks and rollouts.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| GPQA (Diamond) | evaluation | — | — | |
| Terminal-Bench 2.1 | evaluation | — | — | |
| Deep SWE | evaluation | — | — | |
| Terminal-Bench (3.0 / 4.0) | evaluation | — | — | |
| HLE (Humanity's Last Exam) | evaluation | — | — |
Linked Resources
DeepSeek-V4.1 Technical Report
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
deepseek-recipe - protocol-aware prompt encoding toolkit (Rust + Python bindings)
https://github.com/deepseek-ai/deepseek-recipe
Prompt encoding reference (encoding/README.md)
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/encoding/README.md
Minimal inference guide (inference/README.md)
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/inference/README.md
Reproducing DeepSWE v1.1 benchmark results (evaluation/README.md)
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/evaluation/README.md
Pier (datacurve-ai) - integration target for the dsh-minimal agent
https://github.com/datacurve-ai/pier
DeepSeek homepage
https://www.deepseek.com/
DeepSeek Chat
https://chat.deepseek.com/
DeepSeek-V4 HuggingFace collection
https://huggingface.co/collections/deepseek-ai/deepseek-v4
BibTeX citation (DeepSeek-AI, 2026)
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash