DeepSeek-V4.1-Flash

DeepSeek

Parameters

763.0B total / 16.0B active

MoE: total / active

Architecture

MoE

Released

10.09.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16, F32, F8_E4M3, I8 DeepSeek en zh

Input Modalities

text image

Output Modalities

text

Context (native)

1,000,000 tokens

Context (extended)

1,000,000 tokens

Openness Index Score 100.0/100

About

DeepSeek-V4.1-Flash is DeepSeek's MIT-licensed multimodal Mixture-of-Experts model with 552B backbone parameters (763B total) and a 1M-token context window. Maintained by DeepSeek, it natively processes images and text (DeepSeek-ViT vision encoder plus a two-layer MLP projector) and generates text autoregressively, with English and Chinese as primary languages.

Its defining novelty is extreme KV-cache compression for agentic workloads: the global KV cache shrinks to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash and 1/437 of DeepSeek-V1 — while activating only 8B parameters per token during prefill and 16B during decode.

  • Causal Encoder-Decoder (CED): a 40-layer Transformer split into a 20-layer causal encoder and a 20-layer decoder; the decoder's global KV cache is projected from the final encoder hidden states rather than from each decoder layer's own states.
  • Compressed Sparse Attention 2 (CSA2): every attention layer runs in one of three static modes (Full, Reindex, Reuse), sharing main KV and indexer K across layers and reusing Top-K sparse-attention indices; a Hierarchical Sparse Indexer bounds deeper indexer cost independently of context length.
  • FP4 main KV caching (E2M1 format, one E4M3 scale per 16 channels) and SWA Bounded Replay, which reconstructs missing sliding-window-attention KV states by replaying only the most recent n_win tokens — cutting the persistent KV footprint to about 1/8 of DeepSeek-V4-Flash without persisting SWA KV to SSD.
  • Sparse MoE: 1 shared expert plus 384 routed experts per layer with 6 routed experts activated per token; Engram conditional memory (196B parameters, token-based lookup), DSpark speculative decoding (semi-autoregressive drafts with confidence-scheduled verification) and Single-Pass mHC with an efficient Mega-mHC kernel.

Pre-training used a 45T-token multimodal corpus with sparse attention trained at 64K sequence length, extending context to 1M tokens at 34T tokens. Post-training follows the SFT → RL → on-policy distillation (OPD) paradigm, with large-scale automated synthesis of agent tasks and environments; a continuously controllable reasoning_effort setting (integer 1–100) trades inference cost for accuracy.

At maximum reasoning effort the model sets records on DeepSWE v1.1 (74.2) and Terminal-Bench 2.1 (90.6) and leads on CyberGym (88.1), AutomationBench (54.8) and Agent's Last Exam (31.8). Recommended sampling: temperature 1.0, top_p 0.95, context window 1M tokens, max_tokens ≥ 256K.

Training Data Trained from scratch on a 45T-token multimodal corpus; sparse attention trained at 64K sequence length, context extended to 1M tokens at 34T tokens. Post-training: SFT -> RL -> on-policy distillation (OPD) with large-scale automated synthesis of agent tasks and environments.

Benchmark Scores

Benchmark Score Date
GPQA Diamond
stem_reasoning
93.57%
10.09.2026
Humanity's Last Exam
stem_reasoning
65.72%
10.09.2026
MathArena Apex 2025
stem_reasoning
51.04%
10.09.2026
Terminal Bench 2.1
coding_agent
100.00%
10.09.2026
Terminal-Bench 3.0
coding_agent
84.67%
10.09.2026
Terminal-Bench 4.0
coding_agent
100.00%
10.09.2026
DeepSWE 1.1
coding_agent
100.00%
10.09.2026
ProgramBench
coding_agent
25.11%
10.09.2026
NL2Repo-Bench
coding_agent
88.70%
10.09.2026
Cybergym
general_agent
100.00%
10.09.2026
SEC-Bench Pro
general_agent
100.00%
10.09.2026
HLE (with tools)
stem_reasoning
98.73%
10.09.2026
Automation-Bench
general_agent
100.00%
10.09.2026
Agents' Last Exam
general_agent
100.00%
10.09.2026
Chartography
multimodal_agent
100.00%
10.09.2026
BabyVision
vision_language
100.00%
10.09.2026
ZEROBench
vision_language
100.00%
10.09.2026
CodeForces
stem_reasoning
100.00%
10.09.2026

Model tree, quantizations and ecosystem

  • Finetunes: 3 models
  • Quantizations: 23 models (llama.cpp, LM Studio, Jan, Ollama targets)
  • Spaces: 4 spaces using this model
  • Collection: part of the DeepSeek-V4 collection (10 items)
  • Base model for community GGUF variants (e.g. FP8/8-bit/NVFP4 builds)

Citation

@misc{deepseekai2026deepseekv41flash,
      title={DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression},
      author={DeepSeek-AI},
      year={2026},
}

License: MIT

The repository and the model weights are licensed under the MIT License (see the LICENSE file in the repository).

Agentic evaluation setup and scaffold comparison

Code agent benchmarks (Terminal-Bench 2.1/3.0/4.0, DeepSWE v1.1, NL2Repo-Bench, ProgramBench) are evaluated with the Minimal mode of DeepSeek Harness and a 1M-token context window; the mini-SWE harness is used for DeepSWE v1.1 and the Claude Code harness for SEC-Bench Pro. Visual agent benchmarks use the Claude Code harness with 512K context. Scaffold comparison (max effort):

Benchmark Claude Code Codex OpenCode Pi mini-SWE DSH Minimal DSH Standard DSH PTC
DeepSWE v1.1 69.8 65.6 65.5 66.2 74.2 72.6 70.5 67.6
Terminal-Bench 2.1 88.0 84.1 85.0 86.1 90.3 90.6 85.8 85.8

All scaffolds: N=8 samples (DeepSWE) / N=3 (Terminal-Bench 2.1), temperature=1.0, top_p=0.95, 1M-token context limit, max_steps=500, no network access for Terminal-Bench 2.1. The evaluation/ folder reproduces the DeepSWE v1.1 results.

Prompt encoding and local inference

This release does not include a Jinja-format chat template. The encoding/ folder contains a self-contained Python reference implementation (encoding.py) with test cases for multi-turn conversations, tool calling, thinking mode, numeric reasoning effort, mid-conversation system messages, and interleaved image content.

For production use, deepseek-recipe (github.com/deepseek-ai/deepseek-recipe) provides the same prompt format as maintained Rust libraries with Python bindings, covering Messages, Chat Completions, and Responses API requests. The inference/ folder documents weight conversion and local inference.

Recommended sampling parameters

Parameter Value
temperature 1.0
top_p 0.95 or 1.0
context_window 1M tokens
max_tokens >= 256K

Evaluations use temperature=1.0, top_p=0.95.

Continuously controllable reasoning effort

The model supports a continuously controllable reasoning effort setting - an integer from 1 to 100 - that trades inference cost for accuracy. All published instruct results use the maximum effort setting (reasoning_effort=100).

Post-training: SFT -> RL -> on-policy distillation

The post-training recipe follows the standard SFT -> RL -> on-policy distillation (OPD) paradigm without algorithmic modifications. All substantive changes lie in the data pipeline: large-scale automated synthesis of agent tasks and environments with progressive scaling of data, tasks, and rollouts.

Pre-training corpus and context extension

DeepSeek-V4.1-Flash is trained from scratch on a multimodal corpus comprising 45T tokens. Sparse attention is trained at a sequence length of 64K, and context is extended to 1M tokens at the 34T-token mark.

Vision encoder and multimodal pipeline

A vision encoder (DeepSeek-ViT, trained from scratch with 2D-RoPE and 3x3 pixel-unshuffle downsampling) and a two-layer MLP projector convert images into visual embeddings, which are processed jointly with text embeddings from the start of language-model pre-training. Base-model multimodal results: MMMU-Pro 56.5, CVBench 77.9, DocVQA 95.6, RefCOCO-avg 86.0.

MoE design, memory and speculative decoding

  • Sparse MoE: 1 shared expert + 384 routed experts per MoE layer, activating 6 routed experts per token
  • Engram conditional memory: 196B parameters, sparsely accessed via token-based lookup
  • DSPark speculative decoding: semi-autoregressive draft generation with confidence-scheduled verification
  • Single-Pass mHC: revised residual-stream mixing with an efficient Mega-mHC kernel

CSA2 attention and KV-cache compression

Compressed Sparse Attention 2 (CSA2) assigns each attention layer one of three static modes - Full, Reindex, or Reuse - to share main KV and indexer K across layers and reuse Top-K sparse-attention indices. In the decoder, a Hierarchical Sparse Indexer restricts later indexing layers to a candidate pool constructed by the first Full Mode layer, bounding deeper indexer cost independently of context length.

Combined with FP4 main KV caching (E2M1 format, one E4M3 scale per 16 channels), the global KV cache footprint drops to 890 bytes per token - roughly 1/4 of DeepSeek-V4-Flash. SWA Bounded Replay reconstructs missing sliding-window-attention KV states by replaying only the most recent n_win tokens, avoiding the need to persist SWA KV to SSD and reducing the persistent KV cache footprint to about 1/8 of DeepSeek-V4-Flash.

Causal Encoder-Decoder (CED) layout

DeepSeek-V4.1-Flash organizes its 40 Transformer layers as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially improving cost efficiency for input-heavy agentic workloads.

DeepSeek-V4.1-Flash at a glance

  • Multimodal MoE with 552B backbone parameters (763B total), MIT-licensed
  • 1M-token context window, native image + text input, autoregressive text output
  • 8B parameters active per token during prefill, 16B during decode
  • Global KV cache: 890 bytes per token (~1/4 of DeepSeek-V4-Flash, ~1/437 of DeepSeek-V1)
  • Records on DeepSWE v1.1 (74.2) and Terminal-Bench 2.1 (90.6) at max reasoning effort
  • Continuously controllable reasoning effort (integer 1-100)

Architecture

Decoder Block input Embedding vocab 129K · d 5120 Full Attention Sparse Attn 64:1 · dₕ 512 · win 128 ×40 MoE FFN 384 experts · top-6 · +1 shared · dᴻ 2304 MTP Head ×3 speculative layers Final Norm LM Head vocab 129K output
Attention
Sparse Attention (64:1)
MoE
384 experts · top-6 per token
Layers
40
Hidden size
5120
Context
1M tokens
RoPE θ
10K
Parameters
763000M
Active params
16000M

Source: Hugging Face config.json · DeepseekV41ForCausalLM · model repo

Type: Causal Encoder-Decoder (CED) Transformer
Attention: Compressed Sparse Attention 2 (CSA2) with static Full/Reindex/Reuse modes, Hierarchical Sparse Indexer, and sliding-window attention with Bounded Replay
Decoder: 20-layer autoregressive decoder on top of a 20-layer causal encoder (40 layers total)
MoE: yes (384 experts)
Routing: Top-6 routed experts of 384 per token plus 1 always-on shared expert
Layers 40
Total parameters 763B
Context length 1M
Shared experts 1
Vision encoder DeepSeek-ViT trained from scratch with 2D-RoPE and 3x3 pixel-unshuffle downsampling
Active Params Decode 16B
Active Params Prefill 8B
Backbone Params 552B
Conditional Memory Engram (196B parameters, token-based sparse lookup)
Context Extended Tokens 1M
Decoder Layers 20
Encoder Layers 20
Fp4 Main Kv E2M1 with one E4M3 scale per 16 channels
Kv Cache Bytes Per Token 890
Kv Cache Vs V4 Flash ~1/4
Residual Mixing Single-Pass mHC with Mega-mHC kernel
Routed experts/token 6
Sparse Attention Trained Seq Len 66K
Swa Bounded Replay Window n_win most recent tokens replayed to reconstruct SWA KV states
Vision Projector two-layer MLP
Speculative Decoding

DSPark (semi-autoregressive draft generation with confidence-scheduled verification)

Training Pipeline

  1. 1
    pretraining

    Multimodal pre-training from scratch

    45T-token multimodal corpus (text + images); sparse attention trained at a sequence length of 64K.

  2. 2
    cpt

    Context extension to 1M tokens

    Context extended to 1M tokens at the 34T-token mark of pre-training.

  3. 3
    sft

    Supervised fine-tuning

    Standard SFT stage without algorithmic modifications; substantive changes lie in the data pipeline (large-scale automated synthesis of agent tasks and environments).

  4. 4
    rl

    Reinforcement learning

    Standard RL stage of the SFT -> RL -> OPD post-training recipe.

  5. 5
    other

    On-policy distillation (OPD)

    Final consolidation stage; on-policy distillation with progressive scaling of data, tasks and rollouts.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
GPQA (Diamond) evaluation — —
Terminal-Bench 2.1 evaluation — —
Deep SWE evaluation — —
Terminal-Bench (3.0 / 4.0) evaluation — —
HLE (Humanity's Last Exam) evaluation — —

Related Models