Hy4 preview

Tencent

Parameters

770.0B total / 49.0B active

MoE: total / active

Architecture

MoE (78 layers) with Gated DeepSeek Sparse Attention (Gated DSA) + IndexCache cross-layer sparse index reuse, iHC (identity Hyper-Connections) residual streams, 1 native MTP layer (10B total / 0.7B active) for speculative decoding

Released

28.08.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16, F32 Hunyuan en zh

Input Modalities

text

Output Modalities

text

Context (native)

1,000,000 tokens

Context (extended)

1,000,000 tokens

Openness Index Score 100.0/100

About

Hy4 preview (tencent/Hy4-preview, released August 28, 2026 under the Apache 2.0 License) is a new-generation Mixture-of-Experts flagship model from the Tencent Hy Team: 770B total parameters, 49B activated per token, 78 backbone layers (first layer dense FFN, remaining 77 MoE with 256 routed experts and 1 shared expert each; every token activates top-8 routed experts plus the shared expert), plus 1 native MTP layer (10B total / 0.7B activated) for speculative decoding.

Its attention module, inspired by DeepSeek and GLM, employs Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache for cross-layer sparse index reuse (64 attention heads, query compression 2048, KV compression 512, indexer 32 heads/128 dim with top-k 2048); the residual pathway uses iHC (identity Hyper-Connections) expanding inter-layer information flow across 4 residual streams. Hy4 preview scaled on three fronts vs Hy3 - model size, 1M-token context, and substantially larger pre-training and post-training data - the largest generation-over-generation gain the team has measured, putting it at the open-source frontier. Training data was built with Tencent experts (software engineering, game development, finance, security) and co-designed with CodeBuddy and WorkBuddy; in blind side-by-side evaluation 163 internal experts rated it slightly ahead of GLM 5.3 and Kimi K3 on 203 engineering tasks. Reasoning mode defaults to high (deep CoT); recommended sampling temperature 0.9, top_p 1.0.

Training Data Scaled on three fronts vs Hy3: model size, 1M-token context length, and substantially larger pre-training + post-training data. Co-designed with Tencent products (CodeBuddy, WorkBuddy) and internal domain experts (software engineering, game development, finance analysis, security). Reasoning mode defaults to high (deep CoT); recommended sampling temperature=0.9, top_p=1.0.

Benchmark Scores

Benchmark Score Date
SWE-bench Multilingual
coding_agent
92.19%
28.08.2026
SWE-bench Pro
coding_agent
82.12%
28.08.2026
DeepSWE
coding_agent
88.45%
28.08.2026
SWE Atlas - QnA
coding_agent
100.00%
28.08.2026
SWE Atlas - TW
coding_agent
83.05%
28.08.2026
SWE Atlas - RF
coding_agent
87.97%
28.08.2026
SWE-Marathon
coding_agent
63.06%
28.08.2026
Terminal-Bench 2.1 (Terminus-2)
coding_agent
94.99%
28.08.2026
NL2Repo
coding_agent
74.77%
28.08.2026
Cybergym
general_agent
80.36%
28.08.2026
ProgramBench
coding_agent
21.04%
28.08.2026
PostTrainBench
coding_agent
76.79%
28.08.2026
Harbor-Index
coding_agent
58.11%
28.08.2026
Hy-Backend 2.0 (Internal)
coding_agent
38.46%
28.08.2026
Hy-SWE Max Verified (Internal)
coding_agent
72.04%
28.08.2026
Hy-CompanyBench V2 (Internal)
general_agent
75.99%
28.08.2026
WideSearch
general_agent
95.30%
28.08.2026
$OneMillion-Bench
general_capabilities
89.73%
28.08.2026
DRACO
general_agent
51.28%
28.08.2026
Hy-LifeSearch (Internal)
general_agent
42.04%
28.08.2026
Hy-BrowseComp-Pro2 (Internal)
general_agent
64.86%
28.08.2026
OfficeQA Pro
general_agent
98.57%
28.08.2026
MCP-Atlas
general_agent
97.26%
28.08.2026
Toolathlon Verified
general_agent
92.42%
28.08.2026
Apex-Agents
general_agent
87.02%
28.08.2026
SkillsBench Avg5
coding_agent
93.90%
28.08.2026
JobBench
general_agent
86.36%
28.08.2026
WorkSpaceBench
general_agent
11.90%
28.08.2026
Agents' Last Exam
general_agent
57.55%
28.08.2026
GDPVal-AA v2
general_agent
91.64%
28.08.2026
Automation-Bench
general_agent
48.41%
28.08.2026
BankerToolBench
general_agent
81.67%
28.08.2026
E-Bench (Internal)
general_agent
89.10%
28.08.2026
E-Bench-Code (Internal)
coding_agent
77.37%
28.08.2026
Hy-FinAgentBench (Internal)
domain_finance
75.56%
28.08.2026
Hy-FinmodelBench v2 (Internal)
domain_finance
75.94%
28.08.2026
BioMysteryBench
stem_reasoning
90.11%
28.08.2026
HLE (with tools)
stem_reasoning
80.72%
28.08.2026
CritPt (no tools)
stem_reasoning
51.42%
28.08.2026
Humanity's Last Exam
stem_reasoning
78.22%
28.08.2026
SUPERChem
stem_reasoning
57.26%
28.08.2026
ArXivMath
stem_reasoning
53.60%
28.08.2026
HorizonMath (pass@4)
stem_reasoning
74.44%
28.08.2026
MathArena Apex 2025
stem_reasoning
67.36%
28.08.2026
BrokenArXiv
stem_reasoning
54.71%
28.08.2026
GPQA Diamond
stem_reasoning
96.22%
28.08.2026

Hy4 preview - Model Specifications

Backbone specs (excluding MTP layer), from the official model card:

Property Value
Architecture Mixture-of-Experts (MoE)
Total Parameters 770B
Activated Parameters 49B
Layers 78
Hidden Size 6144
Attention Type Gated DSA (Gated DeepSeek Sparse Attention) with IndexCache
Attention Heads 64
Query Compression Dimension 2048
Key-Value Compression Dimension 512
Indexer Heads / Head Dimension 32 / 128
Indexer top-k 2048
Residual Streams 4 (iHC / identity Hyper-Connections)
Routed Experts 256
Shared Experts 1
Activated Routed Experts per Token 8
MoE Intermediate Size 2048
FFN Intermediate Size 18432
Context Length 1M
Vocabulary Size 120832
MTP Layer 10B total / 0.7B activated (speculative decoding)

Recommended sampling: temperature=0.9, top_p=1.0. Reasoning mode defaults to high (deep CoT); disable via chat_template_kwargs: {reasoning_effort: "no_think"}. Served via vLLM (vllm/vllm-openai:hy4-preview, FLASHMLA_SPARSE backend) or SGLang (lmsysorg/sglang:hy4-preview) with MTP/NEXTN speculative decoding (3 draft tokens). Apache 2.0; contact: hunyuan_opensource@tencent.com.

Architecture

Decoder Block input Embedding vocab 121K · d 6144 Full Attention Sparse Attn 64:8 · dₕ 64 ×78 MoE FFN 256 experts · top-8 · +1 shared · dᴻ 2048 MTP Head ×1 speculative layer Final Norm LM Head vocab 121K output
Attention
Sparse Attention (64:8)
MoE
256 experts · top-8 per token
Layers
78
Hidden size
6144
Context
1M tokens
Parameters
770000M
Active params
49000M

Source: Hugging Face config.json · HYV4ForCausalLM · exact layer pattern · model repo

MoE: yes (? experts)

Training Pipeline

  1. 1
    other

    Scaled pre-training + post-training (vs Hy3)

    Hy4 preview was scaled on three fronts: model size, 1M-token context, and training data. Stronger pre-training plus a substantially larger post-training run compound into the largest generation-over-generation gain Tencent measured. Co-designed with Tencent products (CodeBuddy, WorkBuddy) and internal experts from software engineering, game development, finance, and security.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Idavidrein/gpqa (GPQA Diamond) evaluation — human
mercor/apex-agents evaluation — human
harborframework/terminal-bench-2.1 evaluation — human
datacurve/deep-swe evaluation — human
SWE-bench/SWE-bench_Multilingual evaluation — human
ScaleAI/SWE-bench_Pro evaluation — human
benchflow/skillsbench evaluation — human
hkust-nlp/Toolathlon evaluation — human

Trend Analysis

Current

3,516

huggingface

likes

huggingface

followers

huggingface

downloads_all_time

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
huggingface downloads_all_time 3,516 daily 01.09.2026
huggingface followers 11,969 daily 01.09.2026
huggingface likes 381 daily 01.09.2026
huggingface downloads 3,516 daily 01.09.2026

View full metric history →

Related Models