DeepSeek-V4-Flash-Vision-Exp

DeepSeek

Parameters

305.0B total / 13.0B active

MoE: total / active

Architecture

MoE with vision encoder and aligner, DFlash attention, DSpark speculative decoding

Released

08.09.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16, F32, F8_E4M3, I8, I64 DeepSeek en zh

Input Modalities

text image

Output Modalities

text

Context (native)

1,000,000 tokens

Context (extended)

1,000,000 tokens

Openness Index Score 100.0/100

About

DeepSeek-V4-Flash-Vision-Exp (deepseek-ai/DeepSeek-V4-Flash-Vision-Exp) is DeepSeek's first experimental multimodal model in the DeepSeek-V4 family - a 305B-parameter (13B active) sparse Mixture-of-Experts agent model handling text and image input, released September 8, 2026. It builds on the DeepSeek-V4-Flash architecture by incorporating a vision encoder and aligner (32-layer ViT, dim 1024, patch 14, downsample 3, up to 384 tokens per image) and undergoing continued training to unlock visual understanding.

The DeepseekV4 backbone runs DFlash attention (64 heads of dim 512 with 1 KV head, index selection with 64 index heads of dim 128 and top-512, q/o LoRA ranks 1024 in 8 output groups, 3 hash layers) mixed with sparse MoE: 256 routed + 1 shared expert, 6 routed experts per token (expert FFN 2048, fp4 expert dtype, sqrtsoftplus scoring, noaux-tc top-k), 43 layers with per-layer token compression (ratios 4/128, compress RoPE theta 160k), Hyper-Connections (hc_mult 4), a DSpark forward path (block size 5, Markov rank 256, target layers 40-42, noise tokens), 3 Multi-Token Prediction (MTP) layers, native FP8 quantization, sliding window 128, SwiGLU limit 10.0, and a 1,048,576-token context (YaRN factor 16 from 64K).

Compared to DeepSeek-V4-Flash-0731 it achieves substantial improvements on multimodal agent capabilities (ApexBench 36.5 vs 26.2, Agents' Last Exam 27.3 vs 25.2, Chartography 64.3, ZeroBench 35.0) while maintaining comparable text-only agent performance (Terminal Bench 2.1: 83.9). Released under the MIT License.

Training Data First experimental multimodal model in the DeepSeek-V4 family. Builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities. Compared to DeepSeek-V4-Flash-0731, achieves substantial improvements on multimodal agent capabilities while maintaining comparable performance on text-only agent tasks.

Benchmark Scores

Benchmark Score Date
Terminal Bench 2.1
coding_agent
92.57%
08.09.2026
NL2Repo
coding_agent
72.92%
08.09.2026
Cybergym
general_agent
74.09%
08.09.2026
DeepSWE
coding_agent
81.57%
08.09.2026
Toolathlon Verified
general_agent
96.01%
08.09.2026
DSBench-Hard
coding_agent
82.35%
08.09.2026
Automation-Bench
general_agent
33.86%
08.09.2026
ApexBench
multimodal_agent
78.03%
08.09.2026
Chartography
multimodal_agent
64.30
08.09.2026
ZEROBench
vision_language
69.57%
08.09.2026
Agents' Last Exam
general_agent
78.77%
08.09.2026

Model Tree and Spaces

Model tree for deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

Finetunes

10 models

Merges

1 model

Quantizations

25 models

Spaces using deepseek-ai/DeepSeek-V4-Flash-Vision-Exp 3

Collection including deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

[

DeepSeek-V4

Collection

10 items • Updated 1 day ago • 900

](https://huggingface.co/collections/deepseek-ai/deepseek-v4)

License (MIT)

License

This repository is licensed under the MIT License.

Safetensors

Model size

305B params

Tensor type

BF16

·

F32

·

F8_E4M3

·

I8

·

I64

·

How to Use / Deployment

How to Run with vLLM

For example, the command below serves the model with vLLM on a single 4×GB300 node. See the vLLM recipe for detailed instructions and other hardware configurations.

docker run --gpus all \
  vllm/vllm-openai:deepseekv4-flash-vision deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
  --kv-cache-dtype fp8 \
  --block-size 256 \
  --tensor-parallel-size 4 \
  --tool-call-parser deepseek_v4 \
  --enable-auto-tool-choice \
  --reasoning-parser deepseek_v4 \
  --reasoning-config '{"reasoning_parser":"deepseek_v4","reasoning_start_str":"","reasoning_end_str":""}' \
  --speculative-config '{"method":"dspark","model":"deepseek-ai/DeepSeek-V4-Flash-Vision-Exp","num_speculative_tokens":3,"draft_sample_method":"probabilistic","enable_adaptive_verification":true}'

How to Run with SGLang

Enable DSpark with --speculative-algorithm DSPARK and do not set a separate --speculative-draft-model-path as the target and draft weights therefore come from the same checkpoint. See the SGLang cookbook for detailed instructions, benchmarks and other hardwares configurations.

sglang serve \
  --model-path deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
  --tp 4 \
  --speculative-algorithm DSPARK \
  --mem-fraction-static 0.85 \
  --host 0.0.0.0 \
  --port 30000

Minimal Inference (PyTorch reference)

Minimal inference

See inference/README.md for dependency installation, checkpoint conversion, and TXT/JSON inference commands.

Prompt Encoding (OpenAI-style JSON and TXT image notation)

Prompt encoding

See encoding/README.md. Both OpenAI-style JSON content blocks and the compact <image>path</image> TXT notation are supported. The two examples under inference/examples/ encode to identical prompts and token IDs.

Repository Layout (tokenizer, reference inference)

Repository layout

This repository contains the tokenizer, prompt encoding reference, and a minimal PyTorch inference implementation for DeepSeek-V4 Flash Vision. The reference inference covers the vision encoder and aligner, DFlash attention, MoE, Hyper-Connections, and the DSpark forward path.

.
├── encoding/                  # OpenAI-style messages -> model prompt
├── inference/                 # weight conversion and minimal inference
│   └── examples/              # equivalent TXT and JSON vision prompts
├── config.json                # Hugging Face model metadata
├── generation_config.json
├── model.safetensors.index.json
├── tokenizer.json
└── tokenizer_config.json

encoding/ and inference/ deliberately remain separate: prompt formatting does not depend on PyTorch, while inference imports the sibling encoding module with an explicit Python path. No symlinks are required.

The tokenizer files are regular files so that the repository can be uploaded to Hugging Face without relying on local filesystem symlinks. The large model shards are described by model.safetensors.index.json and are not duplicated inside the source checkout used to assemble this repository.

Introduction (vs DeepSeek-V4-Flash-0731, agent benchmarks)

Introduction

We are excited to introduce DeepSeek-V4-Flash-Vision-Exp, our first experimental multimodal model in the DeepSeek-V4 family. It builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities.

Compared to DeepSeek-V4-Flash-0731, DeepSeek-V4-Flash-Vision-Exp achieves substantial improvements on its multimodal agent capabilities, while maintaining comparable performance on text-only agent tasks.

Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Opus-4.8
Text Agent Capabilities
Terminal Bench 2.1 83.9 82.7 85.0
NL2Repo 57.7 54.2 69.7
Cybergym 75.3 76.7 78.3
DeepSWE 59.3 54.4 58.0
Toolathlon-Verified 75.9 70.3 76.2
DSBench-Hard 63.6 59.6 71.7
AutomationBench (Public) 25.7 25.1 27.2
Multimodal Agent Capabilities
ApexBench (Pass@1) 36.5 26.2† 39.4
Agents' Last Exam 27.3 25.2† 25.7
Chartography 64.3 - 65.0
ZeroBench (Pass@5) 35.0 - 34.0

Notes:

  1. For the text agent benchmarks above, DeepSeek models are evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95.
  2. † For ApexBench and Agents' Last Exam, DeepSeek-V4-Flash-0731 ignores the multimodal elements in the input.

Architecture

Decoder Block input Embedding vocab 129K · d 4096 Full Attention Sparse Attn 64:1 · dₕ 512 · win 128 ×43 MoE FFN 256 experts · top-6 · +1 shared · dᴻ 2048 MTP Head ×3 speculative layers Final Norm LM Head vocab 129K output
Attention
Sparse Attention (64:1)
MoE
256 experts · top-6 per token
Layers
43
Hidden size
4096
Context
1M tokens
RoPE θ
10K
Parameters
305000M
Active params
13000M

Source: Hugging Face config.json · DeepseekV4ForCausalLM · model repo

Type: Transformer MoE with Vision Encoder
Attention: DFlash attention
Decoder: autoregressive
MoE: yes (? experts)
Vision Yes
Hyper Connections Yes
MTP Yes
Vision Aligner Yes
Dspark Speculative Decoding

Yes

Training Pipeline

  1. 1
    cpt

    Continued training with visual modules

    Builds on DeepSeek-V4-Flash by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities (card Introduction).

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Vision-language continued training corpus finetune — —

Related Models