Parameters
291.0B total / 13.0B active
MoE: total / active
Architecture
MoE
Released
26.06.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
1,000,000 tokens
Context (extended)
1,000,000 tokens
About
DeepSeek-V4-Flash (deepseek-ai/DeepSeek-V4-Flash) is the preview release of the smaller DeepSeek-V4 MoE model - 284B total parameters with 13B activated per token (card figure; DB stores 291B total) - supporting a one-million-token context, released June 26, 2026 under the MIT License. It supports three reasoning modes: Non-think, Think High, Think Max.
The DeepSeek-V4 architecture combines Compressed Sparse Attention (CSA) with Heavily Compressed Attention (HCA) in a hybrid stack that dramatically improves long-context efficiency - at 1M context the series needs only 27% of single-token inference FLOPs and 10% of the KV cache compared with DeepSeek-V3.2 - and strengthens residual connections with Manifold-Constrained Hyper-Connections (mHC) for stable cross-layer signal propagation. Training uses the Muon optimizer for faster convergence.
Pre-trained on 32T+ diverse high-quality tokens, its post-training runs a two-stage paradigm: independent cultivation of domain-specific experts via SFT and RL with GRPO, then unified model consolidation via on-policy distillation into a single model. DeepSeek-V4-Flash-Max (maximum reasoning effort) achieves reasoning performance comparable to the Pro version given a larger thinking budget, though its smaller scale places it slightly behind on pure knowledge and the most complex agentic workflows.
Training Data Preview version of DeepSeek-V4-Flash. Pre-trained on 32T+ diverse tokens. Post-training: two-stage paradigm with independent domain-specific expert cultivation (SFT + RL with GRPO), then unified model consolidation via on-policy distillation. Supports three reasoning modes: Non-think, Think High, Think Max.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
BrowseComp
general_agent
|
79.57%
|
26.06.2026 |
|
HLE (with tools)
stem_reasoning
|
58.90%
|
26.06.2026 |
|
Vals.ai Financial Agent 1.1 (without web search)
general_agent
|
71.00%
|
04.06.2026 |
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
48.97%
|
04.06.2026 |
|
Vals.ai Financial Agent 1.1 (with web search)
general_agent
|
81.36%
|
04.06.2026 |
|
CritPt (no tools)
stem_reasoning
|
31.55%
|
04.06.2026 |
|
GDPVal
general_agent
|
33.84%
|
04.06.2026 |
|
LiveCodeBench v6
stem_reasoning
|
97.41%
|
26.06.2026 |
|
IMOAnswerBench
stem_reasoning
|
93.47%
|
26.06.2026 |
|
SWE-bench Multilingual
coding_agent
|
80.83%
|
26.06.2026 |
|
OmniScience Accuracy
knowledge
|
73.76%
|
04.06.2026 |
|
ProfBench (Search)
general_agent
|
79.14%
|
04.06.2026 |
|
IMOAnswerBench (with tools)
stem_reasoning
|
77.92%
|
04.06.2026 |
|
OmniScience Non-Hallucination
knowledge
|
2.80
|
04.06.2026 |
|
PinchBench
general_agent
|
100.00%
|
04.06.2026 |
|
TauBench V3 Airline
general_agent
|
52.38%
|
04.06.2026 |
|
IFBench (prompt loose)
instruction_following
|
100.00%
|
04.06.2026 |
|
Apex-Shortlist (with tools)
stem_reasoning
|
86.99%
|
04.06.2026 |
|
Multi-Challenge
instruction_following
|
97.22%
|
04.06.2026 |
|
TauBench V3 Retail
general_agent
|
100.00%
|
04.06.2026 |
|
AA-LCR
long_context
|
78.38%
|
04.06.2026 |
|
RULER (1M)
long_context
|
87.70
|
04.06.2026 |
|
TauBench V3 Telecom
general_agent
|
100.00%
|
04.06.2026 |
|
TauBench V3 Banking
general_agent
|
67.48%
|
04.06.2026 |
|
SciCode (subtask)
stem_reasoning
|
48.06%
|
04.06.2026 |
|
Longbench v2 (≤ 1M)
long_context
|
57.00
|
04.06.2026 |
|
MMLU-ProX
multilingual
|
86.45%
|
04.06.2026 |
|
WMT24++ (en→xx)
multilingual
|
98.20%
|
04.06.2026 |
|
TauBench V3 Average
general_agent
|
100.00%
|
04.06.2026 |
|
Terminal Bench 2.1
coding_agent
|
68.07%
|
13.08.2026 |
|
NL2Repo
coding_agent
|
44.77%
|
13.08.2026 |
|
Cybergym
general_agent
|
38.70
|
13.08.2026 |
|
DeepSWE
coding_agent
|
10.04%
|
13.08.2026 |
|
Toolathlon Verified
general_agent
|
39.92%
|
26.06.2026 |
|
Agents' Last Exam
general_agent
|
24.53%
|
13.08.2026 |
|
Automation-Bench
general_agent
|
10.80
|
13.08.2026 |
|
DSBench-FullStack
coding_agent
|
37.00
|
13.08.2026 |
|
DSBench-Hard
coding_agent
|
25.80
|
13.08.2026 |
|
MMLU-Pro
knowledge
|
84.52%
|
26.06.2026 |
|
GPQA Diamond
stem_reasoning
|
88.28%
|
26.06.2026 |
|
Humanity's Last Exam
stem_reasoning
|
61.93%
|
26.06.2026 |
|
Apex-Shortlist (no tools)
stem_reasoning
|
92.66%
|
26.06.2026 |
|
SWE-bench Verified
coding_agent
|
90.14%
|
26.06.2026 |
|
MMLU
knowledge
|
95.41%
|
26.06.2026 |
|
MMLU-Redux
knowledge
|
73.95%
|
26.06.2026 |
|
MMMLU
multilingual
|
93.45%
|
26.06.2026 |
|
C-Eval
knowledge
|
90.83%
|
26.06.2026 |
|
SuperGPQA
knowledge
|
45.72%
|
26.06.2026 |
|
LongBench v2
long_context
|
44.80%
|
26.06.2026 |
|
MCP-Atlas
general_agent
|
74.97%
|
26.06.2026 |
|
HMMT Feb 26
stem_reasoning
|
96.24%
|
26.06.2026 |
|
Apex-Agents
general_agent
|
75.69%
|
26.06.2026 |
|
Terminal-Bench 2.0
coding_agent
|
68.31%
|
26.06.2026 |
|
SWE-bench Pro
coding_agent
|
65.75%
|
26.06.2026 |
|
GDPVal-AA v2
general_agent
|
76.19%
|
26.06.2026 |
|
CodeForces
stem_reasoning
|
87.53%
|
26.06.2026 |
Model Tree, Spaces and Paper
Model tree for deepseek-ai/DeepSeek-V4-Flash
Adapters
Finetunes
Quantizations
Spaces using deepseek-ai/DeepSeek-V4-Flash 100
Collection including deepseek-ai/DeepSeek-V4-Flash
[
DeepSeek-V4
Collection
10 items • Updated 1 day ago • 900
](https://huggingface.co/collections/deepseek-ai/deepseek-v4)
Paper for deepseek-ai/DeepSeek-V4-Flash
[
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
Paper • 2606.19348 • Published Apr 26 • 43
](https://huggingface.co/papers/2606.19348)
Contact
Contact
If you have any questions, please raise an issue or contact us at service@deepseek.com.
Model size
291B params
Tensor type
BF16
·
I64
·
F32
·
F8_E4M3
·
I8
·
Citation
Citation
@misc{deepseekai2026deepseekv4,
title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
author={DeepSeek-AI},
year={2026},
}
License (MIT)
License
This repository and the model weights are licensed under the MIT License.
How to Run Locally
How to Run Locally
Please refer to the inference folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.
For local deployment, we recommend setting the sampling parameters to temperature = 1.0, top_p = 1.0. For the Think Max reasoning mode, we recommend setting the context window to at least 384K tokens.
Chat Template (reasoning modes)
Chat Template
This release does not include a Jinja-format chat template. Instead, we provide a dedicated encoding folder with Python scripts and test cases demonstrating how to encode messages in OpenAI-compatible format into input strings for the model, and how to parse the model's text output. Please refer to the encoding folder for full documentation.
A brief example:
from encoding_dsv4 import encode_messages, parse_message_from_completion_text
messages = [
{"role": "user", "content": "hello"},
{"role": "assistant", "content": "Hello! I am DeepSeek.", "reasoning_content": "thinking..."},
{"role": "user", "content": "1+1=?"}
]
## messages -> string
prompt = encode_messages(messages, thinking_mode="thinking")
## string -> tokens
import transformers
tokenizer = transformers.AutoTokenizer.from_pretrained("deepseek-ai/DeepSeek-V4-Pro")
tokens = tokenizer.encode(prompt)
Evaluation Results: Instruct Model
Instruct Model
DeepSeek-V4-Pro and DeepSeek-V4-Flash both support three reasoning effort modes:
| Reasoning Mode | Characteristics | Typical Use Cases | Response Format |
|---|---|---|---|
| Non-think | Fast, intuitive responses | Routine daily tasks, low-risk decisions | </think> summary |
| Think High | Conscious logical analysis, slower but more accurate | Complex problem-solving, planning | <think> thinking </think> summary |
| Think Max | Push reasoning to its fullest extent | Exploring the boundary of model reasoning capability | Special system prompt + <think> thinking </think> summary |
DeepSeek-V4-Pro-Max vs Frontier Models
| Benchmark (Metric) | Opus-4.6 Max | GPT-5.4 xHigh | Gemini-3.1-Pro High | K2.6 Thinking | GLM-5.1 Thinking | DS-V4-Pro Max |
|---|---|---|---|---|---|---|
| Knowledge & Reasoning | ||||||
| MMLU-Pro (EM) | 89.1 | 87.5 | 91.0 | 87.1 | 86.0 | 87.5 |
| SimpleQA-Verified (Pass@1) | 46.2 | 45.3 | 75.6 | 36.9 | 38.1 | 57.9 |
| Chinese-SimpleQA (Pass@1) | 76.4 | 76.8 | 85.9 | 75.9 | 75.0 | 84.4 |
| GPQA Diamond (Pass@1) | 91.3 | 93.0 | 94.3 | 90.5 | 86.2 | 90.1 |
| HLE (Pass@1) | 40.0 | 39.8 | 44.4 | 36.4 | 34.7 | 37.7 |
| LiveCodeBench (Pass@1) | 88.8 | - | 91.7 | 89.6 | - | 93.5 |
| Codeforces (Rating) | - | 3168 | 3052 | - | - | 3206 |
| HMMT 2026 Feb (Pass@1) | 96.2 | 97.7 | 94.7 | 92.7 | 89.4 | 95.2 |
| IMOAnswerBench (Pass@1) | 75.3 | 91.4 | 81.0 | 86.0 | 83.8 | 89.8 |
| Apex (Pass@1) | 34.5 | 54.1 | 60.9 | 24.0 | 11.5 | 38.3 |
| Apex Shortlist (Pass@1) | 85.9 | 78.1 | 89.1 | 75.5 | 72.4 | 90.2 |
| Long Context | ||||||
| MRCR 1M (MMR) | 92.9 | - | 76.3 | - | - | 83.5 |
| CorpusQA 1M (ACC) | 71.7 | - | 53.8 | - | - | 62.0 |
| Agentic | ||||||
| Terminal Bench 2.0 (Acc) | 65.4 | 75.1 | 68.5 | 66.7 | 63.5 | 67.9 |
| SWE Verified (Resolved) | 80.8 | - | 80.6 | 80.2 | - | 80.6 |
| SWE Pro (Resolved) | 57.3 | 57.7 | 54.2 | 58.6 | 58.4 | 55.4 |
| SWE Multilingual (Resolved) | 77.5 | - | - | 76.7 | 73.3 | 76.2 |
| BrowseComp (Pass@1) | 83.7 | 82.7 | 85.9 | 83.2 | 79.3 | 83.4 |
| HLE w/ tools (Pass@1) | 53.1 | 52.0 | 51.6 | 54.0 | 50.4 | 48.2 |
| GDPval-AA (Elo) | 1619 | 1674 | 1314 | 1482 | 1535 | 1554 |
| MCPAtlas Public (Pass@1) | 73.8 | 67.2 | 69.2 | 66.6 | 71.8 | 73.6 |
| Toolathlon (Pass@1) | 47.2 | 54.6 | 48.8 | 50.0 | 40.7 | 51.8 |
Comparison across Modes
| Benchmark (Metric) | V4-Flash Non-Think | V4-Flash High | V4-Flash Max | V4-Pro Non-Think | V4-Pro High | V4-Pro Max |
|---|---|---|---|---|---|---|
| Knowledge & Reasoning | ||||||
| MMLU-Pro (EM) | 83.0 | 86.4 | 86.2 | 82.9 | 87.1 | 87.5 |
| SimpleQA-Verified (Pass@1) | 23.1 | 28.9 | 34.1 | 45.0 | 46.2 | 57.9 |
| Chinese-SimpleQA (Pass@1) | 71.5 | 73.2 | 78.9 | 75.8 | 77.7 | 84.4 |
| GPQA Diamond (Pass@1) | 71.2 | 87.4 | 88.1 | 72.9 | 89.1 | 90.1 |
| HLE (Pass@1) | 8.1 | 29.4 | 34.8 | 7.7 | 34.5 | 37.7 |
| LiveCodeBench (Pass@1) | 55.2 | 88.4 | 91.6 | 56.8 | 89.8 | 93.5 |
| Codeforces (Rating) | - | 2816 | 3052 | - | 2919 | 3206 |
| HMMT 2026 Feb (Pass@1) | 40.8 | 91.9 | 94.8 | 31.7 | 94.0 | 95.2 |
| IMOAnswerBench (Pass@1) | 41.9 | 85.1 | 88.4 | 35.3 | 88.0 | 89.8 |
| Apex (Pass@1) | 1.0 | 19.1 | 33.0 | 0.4 | 27.4 | 38.3 |
| Apex Shortlist (Pass@1) | 9.3 | 72.1 | 85.7 | 9.2 | 85.5 | 90.2 |
| Long Context | ||||||
| MRCR 1M (MMR) | 37.5 | 76.9 | 78.7 | 44.7 | 83.3 | 83.5 |
| CorpusQA 1M (ACC) | 15.5 | 59.3 | 60.5 | 35.6 | 56.5 | 62.0 |
| Agentic | ||||||
| Terminal Bench 2.0 (Acc) | 49.1 | 56.6 | 56.9 | 59.1 | 63.3 | 67.9 |
| SWE Verified (Resolved) | 73.7 | 78.6 | 79.0 | 73.6 | 79.4 | 80.6 |
| SWE Pro (Resolved) | 49.1 | 52.3 | 52.6 | 52.1 | 54.4 | 55.4 |
| SWE Multilingual (Resolved) | 69.7 | 70.2 | 73.3 | 69.8 | 74.1 | 76.2 |
| Br |
(table continues in the model card)
Evaluation Results: Base Model
Base Model
| Benchmark (Metric) | # Shots | DeepSeek-V3.2-Base | DeepSeek-V4-Flash-Base | DeepSeek-V4-Pro-Base |
|---|---|---|---|---|
| Architecture | - | MoE | MoE | MoE |
| # Activated Params | - | 37B | 13B | 49B |
| # Total Params | - | 671B | 284B | 1.6T |
| World Knowledge | ||||
| AGIEval (EM) | 0-shot | 80.1 | 82.6 | 83.1 |
| MMLU (EM) | 5-shot | 87.8 | 88.7 | 90.1 |
| MMLU-Redux (EM) | 5-shot | 87.5 | 89.4 | 90.8 |
| MMLU-Pro (EM) | 5-shot | 65.5 | 68.3 | 73.5 |
| MMMLU (EM) | 5-shot | 87.9 | 88.8 | 90.3 |
| C-Eval (EM) | 5-shot | 90.4 | 92.1 | 93.1 |
| CMMLU (EM) | 5-shot | 88.9 | 90.4 | 90.8 |
| MultiLoKo (EM) | 5-shot | 38.7 | 42.2 | 51.1 |
| Simple-QA verified (EM) | 25-shot | 28.3 | 30.1 | 55.2 |
| SuperGPQA (EM) | 5-shot | 45.0 | 46.5 | 53.9 |
| FACTS Parametric (EM) | 25-shot | 27.1 | 33.9 | 62.6 |
| TriviaQA (EM) | 5-shot | 83.3 | 82.8 | 85.6 |
| Language & Reasoning | ||||
| BBH (EM) | 3-shot | 87.6 | 86.9 | 87.5 |
| DROP (F1) | 1-shot | 88.2 | 88.6 | 88.7 |
| HellaSwag (EM) | 0-shot | 86.4 | 85.7 | 88.0 |
| WinoGrande (EM) | 0-shot | 78.9 | 79.5 | 81.5 |
| CLUEWSC (EM) | 5-shot | 83.5 | 82.2 | 85.2 |
| Code & Math | ||||
| BigCodeBench (Pass@1) | 3-shot | 63.9 | 56.8 | 59.2 |
| HumanEval (Pass@1) | 0-shot | 62.8 | 69.5 | 76.8 |
| GSM8K (EM) | 8-shot | 91.1 | 90.8 | 92.6 |
| MATH (EM) | 4-shot | 60.5 | 57.4 | 64.5 |
| MGSM (EM) | 8-shot | 81.3 | 85.7 | 84.4 |
| CMath (EM) | 3-shot | 92.6 | 93.6 | 90.9 |
| Long Context | ||||
| LongBench-V2 (EM) | 1-shot | 40.2 | 44.7 | 51.5 |
Model Downloads (base/instruct/quantizations)
Model Downloads
| Model | #Total Params | #Activated Params | Context Length | Precision | Download |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash-Base | 284B | 13B | 1M | FP8 Mixed | HuggingFace |
| DeepSeek-V4-Flash | 284B | 13B | 1M | FP4 + FP8 Mixed* | HuggingFace |
| DeepSeek-V4-Pro-Base | 1.6T | 49B | 1M | FP8 Mixed | HuggingFace |
| DeepSeek-V4-Pro | 1.6T | 49B | 1M | FP4 + FP8 Mixed* | HuggingFace |
*FP4 + FP8 Mixed: MoE expert parameters use FP4 precision; most other parameters use FP8.
Introduction (DeepSeek-V4 series preview)
Introduction
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.
DeepSeek-V4 series incorporate several key upgrades in architecture and optimization:
- Hybrid Attention Architecture: We design a hybrid attention mechanism combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to dramatically improve long-context efficiency. In the 1M-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2.
- Manifold-Constrained Hyper-Connections (mHC): We incorporate mHC to strengthen conventional residual connections, enhancing stability of signal propagation across layers while preserving model expressivity.
- Muon Optimizer: We employ the Muon optimizer for faster convergence and greater training stability.
We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline. The post-training features a two-stage paradigm: independent cultivation of domain-specific experts (through SFT and RL with GRPO), followed by unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model.
DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, significantly advances the knowledge capabilities of open-source models, firmly establishing itself as the best open-source model available today. It achieves top-tier performance in coding benchmarks and significantly bridges the gap with leading closed-source models on reasoning and agentic tasks. Meanwhile, DeepSeek-V4-Flash-Max achieves comparable reasoning performance to the Pro version when given a larger thinking budget, though its smaller parameter scale naturally places it slightly behind on pure knowledge tasks and the most complex agentic workflows.

Architecture
- Attention
- Sparse Attention (64:1)
- MoE
- 256 experts · top-6 per token
- Layers
- 43
- Hidden size
- 4096
- Context
- 1M tokens
- RoPE θ
- 10K
- Parameters
- 291000M
- Active params
- 13000M
Source: Hugging Face config.json · DeepseekV4ForCausalLM · model repo
Training Pipeline
-
1
pretraining
Pre-training on 32T+ tokens
Pre-trained on more than 32T diverse and high-quality tokens. Uses Muon optimizer for faster convergence and greater training stability. MoE architecture with 284B total params (13B activated).
-
2
sft
Domain-specific expert SFT
Independent cultivation of domain-specific experts through SFT. Part of first stage of two-stage post-training paradigm.
-
3
rl
Domain-specific expert RL (GRPO)
Independent cultivation of domain-specific experts through RL with GRPO. Part of first stage of two-stage post-training paradigm.
-
4
other
On-policy distillation consolidation
Unified model consolidation via on-policy distillation, integrating distinct proficiencies across diverse domains into a single model. Second stage of two-stage post-training paradigm.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| 32T+ token pre-training corpus (diverse, high-quality) | pretraining | — | — | |
| Domain-specific SFT + GRPO RL corpora | rl | — | — |
Linked Resources
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
https://arxiv.org/abs/2606.19348
DeepSeek Homepage
https://www.deepseek.com/
DeepSeek Chat
https://chat.deepseek.com/
DeepSeek-V4 Citation
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash
DeepSeek-V4 Collection
https://huggingface.co/collections/deepseek-ai/deepseek-v4
Trend Analysis
24h Change
+0.1%
7d Change
+0.6%
Current
143,962
downloads
+1.3%
likes
+0.0%
downloads
+0.8%
downloads_all_time
+0.6%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 415,100 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 11,080,030 | daily | 01.09.2026 |
| huggingface | followers | 143,962 | daily | 01.09.2026 |
| huggingface | likes | 2,169 | daily | 01.09.2026 |
| huggingface | downloads | 1,727,171 | daily | 01.09.2026 |
| ollama | downloads | 412,000 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 11,014,089 | daily | 31.08.2026 |
| huggingface | followers | 143,846 | daily | 31.08.2026 |
| huggingface | likes | 2,169 | daily | 31.08.2026 |
| huggingface | downloads | 1,704,533 | daily | 31.08.2026 |
| ollama | downloads | 409,200 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 10,966,688 | daily | 30.08.2026 |
| huggingface | followers | 143,692 | daily | 30.08.2026 |
| huggingface | likes | 2,161 | daily | 30.08.2026 |
| huggingface | downloads | 1,715,319 | daily | 30.08.2026 |
| ollama | downloads | 406,900 pulls | daily | 29.08.2026 |
| huggingface | followers | 143,609 | daily | 29.08.2026 |
| huggingface | likes | 2,157 | daily | 29.08.2026 |
| huggingface | downloads | 1,666,669 | daily | 29.08.2026 |
| ollama | downloads | 404,000 pulls | daily | 28.08.2026 |
| huggingface | followers | 143,516 | daily | 28.08.2026 |
| huggingface | likes | 2,153 | daily | 28.08.2026 |
| huggingface | downloads | 1,636,371 | daily | 28.08.2026 |
| ollama | downloads | 401,600 pulls | daily | 27.08.2026 |
| huggingface | followers | 143,391 | daily | 27.08.2026 |
| huggingface | likes | 2,147 | daily | 27.08.2026 |
| huggingface | downloads | 1,744,766 | daily | 27.08.2026 |
| ollama | downloads | 399,000 pulls | daily | 26.08.2026 |
| huggingface | followers | 143,247 | daily | 26.08.2026 |
| huggingface | likes | 2,143 | daily | 26.08.2026 |
| huggingface | downloads | 1,768,895 | daily | 26.08.2026 |
| ollama | downloads | 396,200 pulls | daily | 25.08.2026 |
| huggingface | followers | 143,114 | daily | 25.08.2026 |
| huggingface | likes | 2,141 | daily | 25.08.2026 |
| huggingface | downloads | 1,769,890 | daily | 25.08.2026 |
| ollama | downloads | 393,600 pulls | daily | 24.08.2026 |
| huggingface | followers | 142,996 | daily | 24.08.2026 |
| huggingface | likes | 2,139 | daily | 24.08.2026 |
| huggingface | downloads | 1,761,668 | daily | 24.08.2026 |