Parameters
305.0B total / 13.0B active
MoE: total / active
Architecture
MoE with vision encoder and aligner, DFlash attention, DSpark speculative decoding
Released
08.09.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
1,000,000 tokens
Context (extended)
1,000,000 tokens
About
DeepSeek-V4-Flash-Vision-Exp (deepseek-ai/DeepSeek-V4-Flash-Vision-Exp) is DeepSeek's first experimental multimodal model in the DeepSeek-V4 family - a 305B-parameter (13B active) sparse Mixture-of-Experts agent model handling text and image input, released September 8, 2026. It builds on the DeepSeek-V4-Flash architecture by incorporating a vision encoder and aligner (32-layer ViT, dim 1024, patch 14, downsample 3, up to 384 tokens per image) and undergoing continued training to unlock visual understanding.
The DeepseekV4 backbone runs DFlash attention (64 heads of dim 512 with 1 KV head, index selection with 64 index heads of dim 128 and top-512, q/o LoRA ranks 1024 in 8 output groups, 3 hash layers) mixed with sparse MoE: 256 routed + 1 shared expert, 6 routed experts per token (expert FFN 2048, fp4 expert dtype, sqrtsoftplus scoring, noaux-tc top-k), 43 layers with per-layer token compression (ratios 4/128, compress RoPE theta 160k), Hyper-Connections (hc_mult 4), a DSpark forward path (block size 5, Markov rank 256, target layers 40-42, noise tokens), 3 Multi-Token Prediction (MTP) layers, native FP8 quantization, sliding window 128, SwiGLU limit 10.0, and a 1,048,576-token context (YaRN factor 16 from 64K).
Compared to DeepSeek-V4-Flash-0731 it achieves substantial improvements on multimodal agent capabilities (ApexBench 36.5 vs 26.2, Agents' Last Exam 27.3 vs 25.2, Chartography 64.3, ZeroBench 35.0) while maintaining comparable text-only agent performance (Terminal Bench 2.1: 83.9). Released under the MIT License.
Training Data First experimental multimodal model in the DeepSeek-V4 family. Builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities. Compared to DeepSeek-V4-Flash-0731, achieves substantial improvements on multimodal agent capabilities while maintaining comparable performance on text-only agent tasks.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Terminal Bench 2.1
coding_agent
|
92.57%
|
08.09.2026 |
|
NL2Repo
coding_agent
|
72.92%
|
08.09.2026 |
|
Cybergym
general_agent
|
74.09%
|
08.09.2026 |
|
DeepSWE
coding_agent
|
81.57%
|
08.09.2026 |
|
Toolathlon Verified
general_agent
|
96.01%
|
08.09.2026 |
|
DSBench-Hard
coding_agent
|
82.35%
|
08.09.2026 |
|
Automation-Bench
general_agent
|
33.86%
|
08.09.2026 |
|
ApexBench
multimodal_agent
|
78.03%
|
08.09.2026 |
|
Chartography
multimodal_agent
|
64.30
|
08.09.2026 |
|
ZEROBench
vision_language
|
69.57%
|
08.09.2026 |
|
Agents' Last Exam
general_agent
|
78.77%
|
08.09.2026 |
Model Tree and Spaces
Model tree for deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
Finetunes
Merges
Quantizations
Spaces using deepseek-ai/DeepSeek-V4-Flash-Vision-Exp 3
Collection including deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
[
DeepSeek-V4
Collection
10 items • Updated 1 day ago • 900
](https://huggingface.co/collections/deepseek-ai/deepseek-v4)
License (MIT)
License
This repository is licensed under the MIT License.
Model size
305B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
I8
·
I64
·
How to Use / Deployment
How to Run with vLLM
For example, the command below serves the model with vLLM on a single 4×GB300 node. See the vLLM recipe for detailed instructions and other hardware configurations.
docker run --gpus all \
vllm/vllm-openai:deepseekv4-flash-vision deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
--kv-cache-dtype fp8 \
--block-size 256 \
--tensor-parallel-size 4 \
--tool-call-parser deepseek_v4 \
--enable-auto-tool-choice \
--reasoning-parser deepseek_v4 \
--reasoning-config '{"reasoning_parser":"deepseek_v4","reasoning_start_str":"","reasoning_end_str":""}' \
--speculative-config '{"method":"dspark","model":"deepseek-ai/DeepSeek-V4-Flash-Vision-Exp","num_speculative_tokens":3,"draft_sample_method":"probabilistic","enable_adaptive_verification":true}'
How to Run with SGLang
Enable DSpark with --speculative-algorithm DSPARK and do not set a separate --speculative-draft-model-path as the target and draft weights therefore come from the same checkpoint. See the SGLang cookbook for detailed instructions, benchmarks and other hardwares configurations.
sglang serve \
--model-path deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
--tp 4 \
--speculative-algorithm DSPARK \
--mem-fraction-static 0.85 \
--host 0.0.0.0 \
--port 30000
Minimal Inference (PyTorch reference)
Minimal inference
See inference/README.md for dependency installation, checkpoint conversion, and TXT/JSON inference commands.
Prompt Encoding (OpenAI-style JSON and TXT image notation)
Prompt encoding
See encoding/README.md. Both OpenAI-style JSON content blocks and the compact <image>path</image> TXT notation are supported. The two examples under inference/examples/ encode to identical prompts and token IDs.
Repository Layout (tokenizer, reference inference)
Repository layout
This repository contains the tokenizer, prompt encoding reference, and a minimal PyTorch inference implementation for DeepSeek-V4 Flash Vision. The reference inference covers the vision encoder and aligner, DFlash attention, MoE, Hyper-Connections, and the DSpark forward path.
.
├── encoding/ # OpenAI-style messages -> model prompt
├── inference/ # weight conversion and minimal inference
│ └── examples/ # equivalent TXT and JSON vision prompts
├── config.json # Hugging Face model metadata
├── generation_config.json
├── model.safetensors.index.json
├── tokenizer.json
└── tokenizer_config.json
encoding/ and inference/ deliberately remain separate: prompt formatting does not depend on PyTorch, while inference imports the sibling encoding module with an explicit Python path. No symlinks are required.
The tokenizer files are regular files so that the repository can be uploaded to Hugging Face without relying on local filesystem symlinks. The large model shards are described by model.safetensors.index.json and are not duplicated inside the source checkout used to assemble this repository.
Introduction (vs DeepSeek-V4-Flash-0731, agent benchmarks)
Introduction
We are excited to introduce DeepSeek-V4-Flash-Vision-Exp, our first experimental multimodal model in the DeepSeek-V4 family. It builds on the DeepSeek-V4-Flash architecture by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities.
Compared to DeepSeek-V4-Flash-0731, DeepSeek-V4-Flash-Vision-Exp achieves substantial improvements on its multimodal agent capabilities, while maintaining comparable performance on text-only agent tasks.
| Benchmark | DeepSeek-V4-Flash-Vision-Exp | DeepSeek-V4-Flash-0731 | Opus-4.8 |
|---|---|---|---|
| Text Agent Capabilities | |||
| Terminal Bench 2.1 | 83.9 | 82.7 | 85.0 |
| NL2Repo | 57.7 | 54.2 | 69.7 |
| Cybergym | 75.3 | 76.7 | 78.3 |
| DeepSWE | 59.3 | 54.4 | 58.0 |
| Toolathlon-Verified | 75.9 | 70.3 | 76.2 |
| DSBench-Hard | 63.6 | 59.6 | 71.7 |
| AutomationBench (Public) | 25.7 | 25.1 | 27.2 |
| Multimodal Agent Capabilities | |||
| ApexBench (Pass@1) | 36.5 | 26.2† | 39.4 |
| Agents' Last Exam | 27.3 | 25.2† | 25.7 |
| Chartography | 64.3 | - | 65.0 |
| ZeroBench (Pass@5) | 35.0 | - | 34.0 |
Notes:
- For the text agent benchmarks above, DeepSeek models are evaluated with the minimal mode of DeepSeek Harness as the agent framework, using the
maxreasoning effort level withtemperature = 1.0, top_p = 0.95. - † For ApexBench and Agents' Last Exam, DeepSeek-V4-Flash-0731 ignores the multimodal elements in the input.
Architecture
- Attention
- Sparse Attention (64:1)
- MoE
- 256 experts · top-6 per token
- Layers
- 43
- Hidden size
- 4096
- Context
- 1M tokens
- RoPE θ
- 10K
- Parameters
- 305000M
- Active params
- 13000M
Source: Hugging Face config.json · DeepseekV4ForCausalLM · model repo
Yes
Training Pipeline
-
1
cpt
Continued training with visual modules
Builds on DeepSeek-V4-Flash by incorporating visual modules and undergoing continued training to unlock visual understanding capabilities (card Introduction).
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Vision-language continued training corpus | finetune | — | — |
Linked Resources
DeepSeek Homepage
https://www.deepseek.com/
DeepSeek Chat
https://chat.deepseek.com/
DeepSeek AI - Hugging Face
https://huggingface.co/deepseek-ai
vLLM Recipe
https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
SGLang Cookbook
https://docs.sglang.io/cookbook/autoregressive/DeepSeek/DeepSeek-V4#hw=b200&variant=flash-vision&quant=fp4&strategy=low-latency&nodes=single