Parameters
295.0B total / 21.0B active
MoE: total / active
Architecture
MoE (80 layers + 1 MTP) with GQA (64 heads/8 KV) + QK-Norm, 192 experts top-8 + 1 shared expert, sigmoid router with expert bias
Released
02.07.2026
License
Apache License 2.0
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
262,144 tokens
About
Hy3 (tencent/Hy3, released July 2, 2026 under the Apache License 2.0) is the Tencent Hy Team's 295B-parameter Mixture-of-Experts model with 21B active parameters plus a 3.8B native MTP layer, introduced after the late-April Hy3 Preview launch with post-training scaled up using feedback from 50+ products. It outperforms similar-size models and rivals flagship open-source models with 2-5x parameters.
Architecture: 80 layers (first dense, then MoE) plus 1 native MTP layer, GQA attention (64 query heads, 8 KV heads, head dim 128) with QK-Norm, 192 routed experts with top-8 activation plus 1 shared expert (sigmoid router with expert bias), hidden size 4096, MoE intermediate size 1536, 256K context, vocabulary 120832, BF16. Hy3 shows solid gains in reasoning, agentic and long-context tasks, and in productivity scenarios (coding, office work, financial modeling, frontend design, game development). In a blind evaluation by 270 experts on real work tasks it scored 2.67/4, outperforming GLM-5.1 at 2.51/4; on SWE-Bench Verified, accuracy variance across scaffoldings (CodeBuddy, Cline, KiloCode) stays within 4%. Product-focused improvements include production-grade tool-call stability, anti-hallucination training (hallucination rate 12.5% -> 5.4%, commonsense errors 25.4% -> 12.7%), and complex multi-turn context retention (issue rate 17.4% -> 7.9%).
Training Data Predecessor of Hy4 preview; scores in the Hy4 preview model card benchmark appendix (re-evaluated by Tencent under the updated harness).
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
SWE-bench Pro
coding_agent
|
72.38%
|
28.08.2026 |
|
DeepSWE
coding_agent
|
38.51%
|
28.08.2026 |
|
SWE Atlas - QnA
coding_agent
|
39.42%
|
28.08.2026 |
|
SWE Atlas - TW
coding_agent
|
48.99%
|
28.08.2026 |
|
SWE Atlas - RF
coding_agent
|
51.35%
|
28.08.2026 |
|
SWE-Marathon
coding_agent
|
8.16%
|
28.08.2026 |
|
NL2Repo
coding_agent
|
54.31%
|
28.08.2026 |
|
Cybergym
general_agent
|
26.52%
|
28.08.2026 |
|
ProgramBench
coding_agent
|
3.00
|
28.08.2026 |
|
PostTrainBench
coding_agent
|
4.78%
|
28.08.2026 |
|
Hy-Backend 2.0 (Internal)
coding_agent
|
26.20
|
28.08.2026 |
|
Hy-SWE Max Verified (Internal)
coding_agent
|
49.00
|
28.08.2026 |
|
WideSearch
general_agent
|
91.39%
|
28.08.2026 |
|
$OneMillion-Bench
general_capabilities
|
36.88%
|
28.08.2026 |
|
Hy-LifeSearch (Internal)
general_agent
|
38.90
|
28.08.2026 |
|
Hy-BrowseComp-Pro2 (Internal)
general_agent
|
57.43%
|
28.08.2026 |
|
OfficeQA Pro
general_agent
|
73.77%
|
28.08.2026 |
|
MCP-Atlas
general_agent
|
85.36%
|
28.08.2026 |
|
Toolathlon Verified
general_agent
|
56.69%
|
28.08.2026 |
|
SkillsBench Avg5
coding_agent
|
81.70%
|
28.08.2026 |
|
JobBench
general_agent
|
27.71%
|
28.08.2026 |
|
Agents' Last Exam
general_agent
|
30.66%
|
28.08.2026 |
|
GDPVal-AA v2
general_agent
|
66.25%
|
28.08.2026 |
|
Automation-Bench
general_agent
|
12.05%
|
28.08.2026 |
|
BankerToolBench
general_agent
|
27.22%
|
28.08.2026 |
|
E-Bench (Internal)
general_agent
|
48.50
|
28.08.2026 |
|
E-Bench-Code (Internal)
coding_agent
|
0.53%
|
28.08.2026 |
|
Hy-FinAgentBench (Internal)
domain_finance
|
69.50
|
28.08.2026 |
|
BioMysteryBench
stem_reasoning
|
54.90
|
28.08.2026 |
|
HLE (with tools)
stem_reasoning
|
73.31%
|
28.08.2026 |
|
CritPt (no tools)
stem_reasoning
|
13.56%
|
28.08.2026 |
|
GPQA Diamond
stem_reasoning
|
93.57%
|
28.08.2026 |
|
Humanity's Last Exam
stem_reasoning
|
61.17%
|
28.08.2026 |
|
ArXivMath
stem_reasoning
|
51.70
|
28.08.2026 |
|
HorizonMath (pass@4)
stem_reasoning
|
3.50
|
28.08.2026 |
|
MathArena Apex 2025
stem_reasoning
|
38.70
|
28.08.2026 |
|
BrokenArXiv
stem_reasoning
|
26.70
|
28.08.2026 |
|
WorkSpaceBench
general_agent
|
58.20
|
28.08.2026 |
|
Hy-FinmodelBench v2 (Internal)
domain_finance
|
28.60
|
28.08.2026 |
|
DRACO
general_agent
|
65.20
|
28.08.2026 |
|
SUPERChem
stem_reasoning
|
52.60
|
28.08.2026 |
|
Harbor-Index
coding_agent
|
15.60
|
28.08.2026 |
|
SWE-bench Multilingual
coding_agent
|
83.79%
|
28.08.2026 |
|
Apex-Agents
general_agent
|
51.93%
|
28.08.2026 |
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
73.45%
|
28.08.2026 |
|
Hy-CompanyBench V2 (Internal)
general_agent
|
29.80
|
28.08.2026 |
Model Tree and Spaces
Model tree for tencent/Hy3
Finetunes
Quantizations
Spaces using tencent/Hy3 25
Collection including tencent/Hy3
[
Hy3
Collection
Hy3 • 2 items • Updated 17 days ago • 22
](https://huggingface.co/collections/tencent/hy3)
Evaluation results
-
mercor/apex-agents · Apex Agents leaderboard
-
harborframework/terminal-bench-2.1 · Terminalbench 2 1 leaderboard
71.7 *
-
SWE-bench/SWE-bench_Multilingual · Swe Bench Multilingual Resolved leaderboard
75.8
-
datacurve/deep-swe · Deep Swe leaderboard
28
-
Idavidrein/gpqa · Diamond [leaderboard](https://huggingface.co/datasets/Idavidrein/gpqa?eval_result=tencent/Hy3&leade
(section continues in the model card)
License
License
Hy3 is released under the Apache License 2.0. See LICENSE for details.
Quantization
Quantization
We provide AngelSlim, a more accessible, comprehensive, and efficient toolkit for large model compression. AngelSlim supports a comprehensive suite of compression tools for large-scale multimodal models, including common quantization algorithms, low-bit quantization, and speculative sampling.
Finetuning and RL Post-training
Finetuning
Hy3 provides a complete model finetuning pipeline. For detailed documentation, please refer to: Finetuning Guide
RL Post-training
Hy3 supports GRPO reinforcement learning training with verl, training on Megatron-LM (model conversion via NVIDIA Megatron-Bridge) with vLLM rollout. For detailed documentation, please refer to: RL Training Guide
Deployment (vLLM, SGLang)
Deployment
Hy3 has 295B parameters in total. To serve it on 8 GPUs, we recommend using H20-3e or other GPUs with larger memory capacity.
For production serving, we recommend using vLLM or SGLang, both of which provide dedicated recipes for Hy3:
-
vLLM - see vLLM recipes
-
SGLang - see SGLang cookbook
vLLM
Build vLLM from source:
uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
git clone https://github.com/vllm-project/vllm.git
cd vllm
uv pip install --editable . --torch-backend=auto
Start the vLLM server with MTP enabled:
## Switch to trtllm backend to work-around mnnvl workspace size issue.
export VLLM_FLASHINFER_ALLREDUCE_BACKEND=trtllm
vllm serve tencent/Hy3 \
--tensor-parallel-size 8 \
--speculative-config.method mtp \
--speculative-config.num_speculative_tokens 2 \
--tool-call-parser hy_v3 \
--reasoning-parser hy_v3 \
--enable-auto-tool-choice \
--port 8000 \
--served-model-name hy3
SGLang
Build SGLang from source:
git clone https://github.com/sgl-project/sglang
cd sglang
pip3 install pip --upgrade
pip3 install "transformers>=5.6.0"
pip3 install -e "python"
Launch SGLang server with MTP enabled:
python3 -m sglang.launch_server \
--model tencent/Hy3 \
--tp-size 8 \
--tool-call-parser hunyuan \
--reasoning-parser hunyuan \
--speculative-num-steps 2 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 3 \
--speculative-algorithm EAGLE \
--port 8000 \
--served-model-name hy3
Quickstart
Quickstart
Deploy Hy3 with vLLM or SGLang first, then call the OpenAI-compatible API:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="hy3",
messages=[
{"role": "user", "content": "Hello! Can you briefly introduce yourself?"},
],
temperature=0.9,
top_p=1.0,
# reasoning_effort: "no_think" (default, direct response), "low", "high" (deep chain-of-thought)
extra_body={"chat_template_kwargs": {"reasoning_effort": "no_think"}},
)
print(response.choices[0].message.content)
Recommended parameters:
temperature=0.9,top_p=1.0.Reasoning mode: Set
reasoning_effortto"high"for complex tasks (math, coding, reasoning) or"no_think"for direct responses.
See the Deployment section below for how to start the API server.
More Reliable Product Experiences
More Reliable Product Experiences
Model usefulness is not fully captured by benchmarks. Based on extensive product feedback, we identified and fixed the following issues, receiving consistently positive feedback from product teams.
Stability of tool calls and output formats: We fixed multiple baseline reliability issues, bringing the model to production-grade standards across tool configurations and output constraints. Tool-call error recovery and overall efficiency improved. Hy3 also generalizes across different agent scaffoldings. On SWE-Bench Verified, accuracy variance across scaffoldings like CodeBuddy, Cline, and KiloCode remains within 4%.
Knowledge and anti-hallucination: Guided by the ideal of "answer when grounded, state when evidence is missing, do not conflate sources or fabricate data," we implemented fine-grained data cleaning and training constraints. In internal evaluations based on real-world scenarios, Hy3's hallucination rate dropped from 12.5% to 5.4%, and commonsense error rates fell from 25.4% to 12.7%. These improvements materially reduce fact conflation, fabrication, and logical contradiction.
Complex context retention and multi-turn intent tracking: Through joint optimization of SFT and RL, Hy3 improved on operational pain points like coreference resolution, ellipsis recovery, and multi-turn constraint inheritance. On internal comprehensive multi-turn tests, the issue rate dropped from 17.4% to 7.9%. Hy3 also improved markedly on long-dialogue evals like MRCR. Its outputs are more concise while ensuring complex intents do not decay or drift over long-horizon interactions.
Benchmark Appendix

Stronger Agent Capabilities
Stronger Agent Capabilities
Building on Hy3 Preview, we further improved the quality and diversity of post-training data while scaling up RL training. Hy3 shows solid gains across reasoning, agentic, and long-context tasks, competitive with much larger flagship models.

In productivity scenarios such as coding, office work, financial modeling, frontend design, and game development, Hy3 has made remarkable progress and can now serve as a reliable, cost-effective model option.
We don't think public benchmark scores tell the full story. So we ran a blind evaluation with 270 experts using tasks from their work, and Hy3 scored 2.67/4, outperforming GLM-5.1 at 2.51/4. The advantage was most substantial in frontend development, data & storage, and CI/CD tasks.
Model Introduction (architecture table)
Model Introduction
Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data. Today, we introduce Hy3, which outperforms similar-size models and rivals flagship open-source models with 2-5x parameters. It also shows significant gains in utility across various products and productivity tasks.
| Property | Value |
|---|---|
| Architecture | Mixture-of-Experts (MoE) |
| Total Parameters | 295B |
| Activated Parameters | 21B |
| MTP Layer Parameters | 3.8B |
| Number of Layers (excluding MTP layer) | 80 |
| Number of MTP Layers | 1 |
| Attention Heads | 64 (GQA, 8 KV heads, head dim 128) |
| Hidden Size | 4096 |
| Intermediate Size | 13312 |
| Context Length | 256K |
| Vocabulary Size | 120832 |
| Number of Experts | 192 experts, top-8 activated |
| Supported Precisions | BF16 |
Architecture
- Attention
- Grouped Query Attention (64:8)
- MoE
- 192 experts · top-8 per token
- Layers
- 80
- Hidden size
- 4096
- Context
- 262K tokens
- Parameters
- 295000M
- Active params
- 21000M
Source: Hugging Face config.json · HYV3ForCausalLM · model repo
Training Pipeline
-
2
sft
SFT stage
Post-training SFT with higher quality and more diverse data scaled from 50+ products feedback; joint SFT+RL optimization for multi-turn context retention
-
3
rl
RL stage
Scaled-up RL (GRPO with verl, Megatron-LM training, vLLM rollout) across reasoning, agentic and long-context tasks