Hy3

Tencent

Parameters

295.0B total / 21.0B active

MoE: total / active

Architecture

MoE (80 layers + 1 MTP) with GQA (64 heads/8 KV) + QK-Norm, 192 experts top-8 + 1 shared expert, sigmoid router with expert bias

Released

02.07.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16 Hunyuan en zh

Input Modalities

text

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

262,144 tokens

Openness Index Score 100.0/100

About

Hy3 (tencent/Hy3, released July 2, 2026 under the Apache License 2.0) is the Tencent Hy Team's 295B-parameter Mixture-of-Experts model with 21B active parameters plus a 3.8B native MTP layer, introduced after the late-April Hy3 Preview launch with post-training scaled up using feedback from 50+ products. It outperforms similar-size models and rivals flagship open-source models with 2-5x parameters.

Architecture: 80 layers (first dense, then MoE) plus 1 native MTP layer, GQA attention (64 query heads, 8 KV heads, head dim 128) with QK-Norm, 192 routed experts with top-8 activation plus 1 shared expert (sigmoid router with expert bias), hidden size 4096, MoE intermediate size 1536, 256K context, vocabulary 120832, BF16. Hy3 shows solid gains in reasoning, agentic and long-context tasks, and in productivity scenarios (coding, office work, financial modeling, frontend design, game development). In a blind evaluation by 270 experts on real work tasks it scored 2.67/4, outperforming GLM-5.1 at 2.51/4; on SWE-Bench Verified, accuracy variance across scaffoldings (CodeBuddy, Cline, KiloCode) stays within 4%. Product-focused improvements include production-grade tool-call stability, anti-hallucination training (hallucination rate 12.5% -> 5.4%, commonsense errors 25.4% -> 12.7%), and complex multi-turn context retention (issue rate 17.4% -> 7.9%).

Training Data Predecessor of Hy4 preview; scores in the Hy4 preview model card benchmark appendix (re-evaluated by Tencent under the updated harness).

Benchmark Scores

Benchmark Score Date
SWE-bench Pro
coding_agent
72.38%
28.08.2026
DeepSWE
coding_agent
38.51%
28.08.2026
SWE Atlas - QnA
coding_agent
39.42%
28.08.2026
SWE Atlas - TW
coding_agent
48.99%
28.08.2026
SWE Atlas - RF
coding_agent
51.35%
28.08.2026
SWE-Marathon
coding_agent
8.16%
28.08.2026
NL2Repo
coding_agent
54.31%
28.08.2026
Cybergym
general_agent
26.52%
28.08.2026
ProgramBench
coding_agent
3.00
28.08.2026
PostTrainBench
coding_agent
4.78%
28.08.2026
Hy-Backend 2.0 (Internal)
coding_agent
26.20
28.08.2026
Hy-SWE Max Verified (Internal)
coding_agent
49.00
28.08.2026
WideSearch
general_agent
91.39%
28.08.2026
$OneMillion-Bench
general_capabilities
36.88%
28.08.2026
Hy-LifeSearch (Internal)
general_agent
38.90
28.08.2026
Hy-BrowseComp-Pro2 (Internal)
general_agent
57.43%
28.08.2026
OfficeQA Pro
general_agent
73.77%
28.08.2026
MCP-Atlas
general_agent
85.36%
28.08.2026
Toolathlon Verified
general_agent
56.69%
28.08.2026
SkillsBench Avg5
coding_agent
81.70%
28.08.2026
JobBench
general_agent
27.71%
28.08.2026
Agents' Last Exam
general_agent
30.66%
28.08.2026
GDPVal-AA v2
general_agent
66.25%
28.08.2026
Automation-Bench
general_agent
12.05%
28.08.2026
BankerToolBench
general_agent
27.22%
28.08.2026
E-Bench (Internal)
general_agent
48.50
28.08.2026
E-Bench-Code (Internal)
coding_agent
0.53%
28.08.2026
Hy-FinAgentBench (Internal)
domain_finance
69.50
28.08.2026
BioMysteryBench
stem_reasoning
54.90
28.08.2026
HLE (with tools)
stem_reasoning
73.31%
28.08.2026
CritPt (no tools)
stem_reasoning
13.56%
28.08.2026
GPQA Diamond
stem_reasoning
93.57%
28.08.2026
Humanity's Last Exam
stem_reasoning
61.17%
28.08.2026
ArXivMath
stem_reasoning
51.70
28.08.2026
HorizonMath (pass@4)
stem_reasoning
3.50
28.08.2026
MathArena Apex 2025
stem_reasoning
38.70
28.08.2026
BrokenArXiv
stem_reasoning
26.70
28.08.2026
WorkSpaceBench
general_agent
58.20
28.08.2026
Hy-FinmodelBench v2 (Internal)
domain_finance
28.60
28.08.2026
DRACO
general_agent
65.20
28.08.2026
SUPERChem
stem_reasoning
52.60
28.08.2026
Harbor-Index
coding_agent
15.60
28.08.2026
SWE-bench Multilingual
coding_agent
83.79%
28.08.2026
Apex-Agents
general_agent
51.93%
28.08.2026
Terminal-Bench 2.1 (Terminus-2)
coding_agent
73.45%
28.08.2026
Hy-CompanyBench V2 (Internal)
general_agent
29.80
28.08.2026

Model Tree and Spaces

Model tree for tencent/Hy3

Finetunes

10 models

Quantizations

71 models

Spaces using tencent/Hy3 25

Collection including tencent/Hy3

[

Hy3

Collection

Hy3 • 2 items • Updated 17 days ago • 22

](https://huggingface.co/collections/tencent/hy3)

Evaluation results

(section continues in the model card)

License

License

Hy3 is released under the Apache License 2.0. See LICENSE for details.

Quantization

Quantization

We provide AngelSlim, a more accessible, comprehensive, and efficient toolkit for large model compression. AngelSlim supports a comprehensive suite of compression tools for large-scale multimodal models, including common quantization algorithms, low-bit quantization, and speculative sampling.

Finetuning and RL Post-training

Finetuning

Hy3 provides a complete model finetuning pipeline. For detailed documentation, please refer to: Finetuning Guide

RL Post-training

Hy3 supports GRPO reinforcement learning training with verl, training on Megatron-LM (model conversion via NVIDIA Megatron-Bridge) with vLLM rollout. For detailed documentation, please refer to: RL Training Guide

Deployment (vLLM, SGLang)

Deployment

Hy3 has 295B parameters in total. To serve it on 8 GPUs, we recommend using H20-3e or other GPUs with larger memory capacity.

For production serving, we recommend using vLLM or SGLang, both of which provide dedicated recipes for Hy3:

vLLM

Build vLLM from source:

uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
git clone https://github.com/vllm-project/vllm.git
cd vllm
uv pip install --editable . --torch-backend=auto

Start the vLLM server with MTP enabled:

## Switch to trtllm backend to work-around mnnvl workspace size issue.
export VLLM_FLASHINFER_ALLREDUCE_BACKEND=trtllm
vllm serve tencent/Hy3 \
  --tensor-parallel-size 8 \
  --speculative-config.method mtp \
  --speculative-config.num_speculative_tokens 2 \
  --tool-call-parser hy_v3 \
  --reasoning-parser hy_v3 \
  --enable-auto-tool-choice \
  --port 8000 \
  --served-model-name hy3

SGLang

Build SGLang from source:

git clone https://github.com/sgl-project/sglang
cd sglang
pip3 install pip --upgrade
pip3 install "transformers>=5.6.0"
pip3 install -e "python"

Launch SGLang server with MTP enabled:

python3 -m sglang.launch_server \
  --model tencent/Hy3 \
  --tp-size 8 \
  --tool-call-parser hunyuan \
  --reasoning-parser hunyuan \
  --speculative-num-steps 2 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 3 \
  --speculative-algorithm EAGLE \
  --port 8000 \
  --served-model-name hy3

Quickstart

Quickstart

Deploy Hy3 with vLLM or SGLang first, then call the OpenAI-compatible API:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="hy3",
    messages=[
        {"role": "user", "content": "Hello! Can you briefly introduce yourself?"},
    ],
    temperature=0.9,
    top_p=1.0,
    # reasoning_effort: "no_think" (default, direct response), "low", "high" (deep chain-of-thought)
    extra_body={"chat_template_kwargs": {"reasoning_effort": "no_think"}},
)
print(response.choices[0].message.content)

Recommended parameters: temperature=0.9, top_p=1.0.

Reasoning mode: Set reasoning_effort to "high" for complex tasks (math, coding, reasoning) or "no_think" for direct responses.

See the Deployment section below for how to start the API server.

Model Links (HF, ModelScope, GitCode, CNB)

Model Links

Model Name Description Hugging Face ModelScope GitCode CNB
Hy3 Instruct model 🤗 Model Model Model Model
Hy3-FP8 FP8 quantized instruct model 🤗 Model Model Model Model

More Reliable Product Experiences

More Reliable Product Experiences

Model usefulness is not fully captured by benchmarks. Based on extensive product feedback, we identified and fixed the following issues, receiving consistently positive feedback from product teams.

Stability of tool calls and output formats: We fixed multiple baseline reliability issues, bringing the model to production-grade standards across tool configurations and output constraints. Tool-call error recovery and overall efficiency improved. Hy3 also generalizes across different agent scaffoldings. On SWE-Bench Verified, accuracy variance across scaffoldings like CodeBuddy, Cline, and KiloCode remains within 4%.

Knowledge and anti-hallucination: Guided by the ideal of "answer when grounded, state when evidence is missing, do not conflate sources or fabricate data," we implemented fine-grained data cleaning and training constraints. In internal evaluations based on real-world scenarios, Hy3's hallucination rate dropped from 12.5% to 5.4%, and commonsense error rates fell from 25.4% to 12.7%. These improvements materially reduce fact conflation, fabrication, and logical contradiction.

Complex context retention and multi-turn intent tracking: Through joint optimization of SFT and RL, Hy3 improved on operational pain points like coreference resolution, ellipsis recovery, and multi-turn constraint inheritance. On internal comprehensive multi-turn tests, the issue rate dropped from 17.4% to 7.9%. Hy3 also improved markedly on long-dialogue evals like MRCR. Its outputs are more concise while ensuring complex intents do not decay or drift over long-horizon interactions.

Benchmark Appendix

Stronger Agent Capabilities

Stronger Agent Capabilities

Building on Hy3 Preview, we further improved the quality and diversity of post-training data while scaling up RL training. Hy3 shows solid gains across reasoning, agentic, and long-context tasks, competitive with much larger flagship models.

In productivity scenarios such as coding, office work, financial modeling, frontend design, and game development, Hy3 has made remarkable progress and can now serve as a reliable, cost-effective model option.

We don't think public benchmark scores tell the full story. So we ran a blind evaluation with 270 experts using tasks from their work, and Hy3 scored 2.67/4, outperforming GLM-5.1 at 2.51/4. The advantage was most substantial in frontend development, data & storage, and CI/CD tasks.

Model Introduction (architecture table)

Model Introduction

Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tencent Hy Team. Following the Hy3 Preview launch in late April, we gathered feedback from 50+ products and scaled up post-training with higher quality data. Today, we introduce Hy3, which outperforms similar-size models and rivals flagship open-source models with 2-5x parameters. It also shows significant gains in utility across various products and productivity tasks.

Property Value
Architecture Mixture-of-Experts (MoE)
Total Parameters 295B
Activated Parameters 21B
MTP Layer Parameters 3.8B
Number of Layers (excluding MTP layer) 80
Number of MTP Layers 1
Attention Heads 64 (GQA, 8 KV heads, head dim 128)
Hidden Size 4096
Intermediate Size 13312
Context Length 256K
Vocabulary Size 120832
Number of Experts 192 experts, top-8 activated
Supported Precisions BF16

Architecture

Decoder Block input Embedding vocab 121K · d 4096 Full Attention GQA 64:8 · dₕ 128 ×80 MoE FFN 192 experts · top-8 · +1 shared · dᴻ 1536 MTP Head ×1 speculative layer Final Norm LM Head vocab 121K output
Attention
Grouped Query Attention (64:8)
MoE
192 experts · top-8 per token
Layers
80
Hidden size
4096
Context
262K tokens
Parameters
295000M
Active params
21000M

Source: Hugging Face config.json · HYV3ForCausalLM · model repo

Type: Decoder-only MoE Transformer
Attention: Grouped-Query Attention (GQA)
Decoder: Autoregressive (80 layers + 1 native MTP)
MoE: yes (192 experts)
Routing: Top-8 of 192 routed experts + 1 shared expert; sigmoid router with expert bias
Layers 80
Context length 262K
Extended context 262K
Experts 192
Experts per token 8
Attention heads 64
KV heads 8
Hidden size 4096
Expert FFN dim 1536
Vision No
MTP Yes
RoPE dim 128

Training Pipeline

  1. 2
    sft

    SFT stage

    Post-training SFT with higher quality and more diverse data scaled from 50+ products feedback; joint SFT+RL optimization for multi-turn context retention

  2. 3
    rl

    RL stage

    Scaled-up RL (GRPO with verl, Megatron-LM training, vLLM rollout) across reasoning, agentic and long-context tasks

Related Models