Parameters
770.0B total / 49.0B active
MoE: total / active
Architecture
MoE (78 layers) with Gated DeepSeek Sparse Attention (Gated DSA) + IndexCache cross-layer sparse index reuse, iHC (identity Hyper-Connections) residual streams, 1 native MTP layer (10B total / 0.7B active) for speculative decoding
Released
28.08.2026
License
Apache License 2.0
Input Modalities
Output Modalities
Context (native)
1,000,000 tokens
Context (extended)
1,000,000 tokens
About
Hy4 preview (tencent/Hy4-preview, released August 28, 2026 under the Apache 2.0 License) is a new-generation Mixture-of-Experts flagship model from the Tencent Hy Team: 770B total parameters, 49B activated per token, 78 backbone layers (first layer dense FFN, remaining 77 MoE with 256 routed experts and 1 shared expert each; every token activates top-8 routed experts plus the shared expert), plus 1 native MTP layer (10B total / 0.7B activated) for speculative decoding.
Its attention module, inspired by DeepSeek and GLM, employs Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache for cross-layer sparse index reuse (64 attention heads, query compression 2048, KV compression 512, indexer 32 heads/128 dim with top-k 2048); the residual pathway uses iHC (identity Hyper-Connections) expanding inter-layer information flow across 4 residual streams. Hy4 preview scaled on three fronts vs Hy3 - model size, 1M-token context, and substantially larger pre-training and post-training data - the largest generation-over-generation gain the team has measured, putting it at the open-source frontier. Training data was built with Tencent experts (software engineering, game development, finance, security) and co-designed with CodeBuddy and WorkBuddy; in blind side-by-side evaluation 163 internal experts rated it slightly ahead of GLM 5.3 and Kimi K3 on 203 engineering tasks. Reasoning mode defaults to high (deep CoT); recommended sampling temperature 0.9, top_p 1.0.
Training Data Scaled on three fronts vs Hy3: model size, 1M-token context length, and substantially larger pre-training + post-training data. Co-designed with Tencent products (CodeBuddy, WorkBuddy) and internal domain experts (software engineering, game development, finance analysis, security). Reasoning mode defaults to high (deep CoT); recommended sampling temperature=0.9, top_p=1.0.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
SWE-bench Multilingual
coding_agent
|
92.19%
|
28.08.2026 |
|
SWE-bench Pro
coding_agent
|
82.12%
|
28.08.2026 |
|
DeepSWE
coding_agent
|
88.45%
|
28.08.2026 |
|
SWE Atlas - QnA
coding_agent
|
100.00%
|
28.08.2026 |
|
SWE Atlas - TW
coding_agent
|
83.05%
|
28.08.2026 |
|
SWE Atlas - RF
coding_agent
|
87.97%
|
28.08.2026 |
|
SWE-Marathon
coding_agent
|
63.06%
|
28.08.2026 |
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
94.99%
|
28.08.2026 |
|
NL2Repo
coding_agent
|
74.77%
|
28.08.2026 |
|
Cybergym
general_agent
|
80.36%
|
28.08.2026 |
|
ProgramBench
coding_agent
|
21.04%
|
28.08.2026 |
|
PostTrainBench
coding_agent
|
76.79%
|
28.08.2026 |
|
Harbor-Index
coding_agent
|
58.11%
|
28.08.2026 |
|
Hy-Backend 2.0 (Internal)
coding_agent
|
38.46%
|
28.08.2026 |
|
Hy-SWE Max Verified (Internal)
coding_agent
|
72.04%
|
28.08.2026 |
|
Hy-CompanyBench V2 (Internal)
general_agent
|
75.99%
|
28.08.2026 |
|
WideSearch
general_agent
|
95.30%
|
28.08.2026 |
|
$OneMillion-Bench
general_capabilities
|
89.73%
|
28.08.2026 |
|
DRACO
general_agent
|
51.28%
|
28.08.2026 |
|
Hy-LifeSearch (Internal)
general_agent
|
42.04%
|
28.08.2026 |
|
Hy-BrowseComp-Pro2 (Internal)
general_agent
|
64.86%
|
28.08.2026 |
|
OfficeQA Pro
general_agent
|
98.57%
|
28.08.2026 |
|
MCP-Atlas
general_agent
|
97.26%
|
28.08.2026 |
|
Toolathlon Verified
general_agent
|
92.42%
|
28.08.2026 |
|
Apex-Agents
general_agent
|
87.02%
|
28.08.2026 |
|
SkillsBench Avg5
coding_agent
|
93.90%
|
28.08.2026 |
|
JobBench
general_agent
|
86.36%
|
28.08.2026 |
|
WorkSpaceBench
general_agent
|
11.90%
|
28.08.2026 |
|
Agents' Last Exam
general_agent
|
57.55%
|
28.08.2026 |
|
GDPVal-AA v2
general_agent
|
91.64%
|
28.08.2026 |
|
Automation-Bench
general_agent
|
48.41%
|
28.08.2026 |
|
BankerToolBench
general_agent
|
81.67%
|
28.08.2026 |
|
E-Bench (Internal)
general_agent
|
89.10%
|
28.08.2026 |
|
E-Bench-Code (Internal)
coding_agent
|
77.37%
|
28.08.2026 |
|
Hy-FinAgentBench (Internal)
domain_finance
|
75.56%
|
28.08.2026 |
|
Hy-FinmodelBench v2 (Internal)
domain_finance
|
75.94%
|
28.08.2026 |
|
BioMysteryBench
stem_reasoning
|
90.11%
|
28.08.2026 |
|
HLE (with tools)
stem_reasoning
|
80.72%
|
28.08.2026 |
|
CritPt (no tools)
stem_reasoning
|
51.42%
|
28.08.2026 |
|
Humanity's Last Exam
stem_reasoning
|
78.22%
|
28.08.2026 |
|
SUPERChem
stem_reasoning
|
57.26%
|
28.08.2026 |
|
ArXivMath
stem_reasoning
|
53.60%
|
28.08.2026 |
|
HorizonMath (pass@4)
stem_reasoning
|
74.44%
|
28.08.2026 |
|
MathArena Apex 2025
stem_reasoning
|
67.36%
|
28.08.2026 |
|
BrokenArXiv
stem_reasoning
|
54.71%
|
28.08.2026 |
|
GPQA Diamond
stem_reasoning
|
96.22%
|
28.08.2026 |
Hy4 preview - Model Specifications
Backbone specs (excluding MTP layer), from the official model card:
| Property | Value |
|---|---|
| Architecture | Mixture-of-Experts (MoE) |
| Total Parameters | 770B |
| Activated Parameters | 49B |
| Layers | 78 |
| Hidden Size | 6144 |
| Attention Type | Gated DSA (Gated DeepSeek Sparse Attention) with IndexCache |
| Attention Heads | 64 |
| Query Compression Dimension | 2048 |
| Key-Value Compression Dimension | 512 |
| Indexer Heads / Head Dimension | 32 / 128 |
| Indexer top-k | 2048 |
| Residual Streams | 4 (iHC / identity Hyper-Connections) |
| Routed Experts | 256 |
| Shared Experts | 1 |
| Activated Routed Experts per Token | 8 |
| MoE Intermediate Size | 2048 |
| FFN Intermediate Size | 18432 |
| Context Length | 1M |
| Vocabulary Size | 120832 |
| MTP Layer | 10B total / 0.7B activated (speculative decoding) |
Recommended sampling: temperature=0.9, top_p=1.0. Reasoning mode defaults to high (deep CoT); disable via chat_template_kwargs: {reasoning_effort: "no_think"}. Served via vLLM (vllm/vllm-openai:hy4-preview, FLASHMLA_SPARSE backend) or SGLang (lmsysorg/sglang:hy4-preview) with MTP/NEXTN speculative decoding (3 draft tokens). Apache 2.0; contact: hunyuan_opensource@tencent.com.
Architecture
- Attention
- Sparse Attention (64:8)
- MoE
- 256 experts · top-8 per token
- Layers
- 78
- Hidden size
- 6144
- Context
- 1M tokens
- Parameters
- 770000M
- Active params
- 49000M
Source: Hugging Face config.json · HYV4ForCausalLM · exact layer pattern · model repo
Training Pipeline
-
1
other
Scaled pre-training + post-training (vs Hy3)
Hy4 preview was scaled on three fronts: model size, 1M-token context, and training data. Stronger pre-training plus a substantially larger post-training run compound into the largest generation-over-generation gain Tencent measured. Co-designed with Tencent products (CodeBuddy, WorkBuddy) and internal experts from software engineering, game development, finance, and security.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Idavidrein/gpqa (GPQA Diamond) | evaluation | — | human | |
| mercor/apex-agents | evaluation | — | human | |
| harborframework/terminal-bench-2.1 | evaluation | — | human | |
| datacurve/deep-swe | evaluation | — | human | |
| SWE-bench/SWE-bench_Multilingual | evaluation | — | human | |
| ScaleAI/SWE-bench_Pro | evaluation | — | human | |
| benchflow/skillsbench | evaluation | — | human | |
| hkust-nlp/Toolathlon | evaluation | — | human |
Linked Resources
Official Website (Tencent AI Studio)
https://aistudio.tencent.com/
GitHub Repository Tencent-Hunyuan/Hy4-preview
https://github.com/Tencent-Hunyuan/Hy4-preview
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
https://arxiv.org/abs/2603.12201
DeepSeek-V3.2: Gated DeepSeek Sparse Attention (DSA) reference
https://arxiv.org/abs/2512.02556
iHC (identity Hyper-Connections) - Zhihu article
https://zhuanlan.zhihu.com/p/2010852389670908320
Hy4 preview collection (Hugging Face)
https://huggingface.co/collections/tencent/hy4-preview
Hy4 preview on ModelScope
https://modelscope.cn/models/Tencent-Hunyuan/Hy4-preview
Hy4 preview on GitCode
https://ai.gitcode.com/tencent_hunyuan/Hy4-preview
Hy4 preview on CNB
https://cnb.cool/ai-models/tencent/Hy4-preview
FP8 quantized variant: tencent/Hy4-preview-FP8
https://huggingface.co/tencent/Hy4-preview-FP8
vLLM deployment recipe
https://recipes.vllm.ai/tencent/Hy4-preview
SGLang deployment cookbook
https://lmsysorg.mintlify.app/cookbook/autoregressive/Tencent/Hy4-Preview
Finetuning guide
https://huggingface.co/tencent/Hy4-preview/blob/main/finetune/README.md
AngelSlim quantization toolkit (Tencent)
https://github.com/tencent/AngelSlim
Benchmark appendix image (source of scores)
https://huggingface.co/tencent/Hy4-preview/resolve/main/assets/benchmark-appendix.jpg
Trend Analysis
Current
3,516
likes
followers
downloads_all_time
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| huggingface | downloads_all_time | 3,516 | daily | 01.09.2026 |
| huggingface | followers | 11,969 | daily | 01.09.2026 |
| huggingface | likes | 381 | daily | 01.09.2026 |
| huggingface | downloads | 3,516 | daily | 01.09.2026 |