Xing4.0-29B-A4B

China Telecom AI

Parameters

31.2B total / 4.0B active

MoE: total / active

Architecture

Sparse MoE Transformer with mHC + MLA + MTP

Released

16.09.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16,F32 Xing4.0 English Chinese

Input Modalities

text

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

524,288 tokens

Openness Index Score 100.0/100

About

Xing4.0-29B-A4B (XingChen-AGI/Xing4.0-29B-A4B) is the next-generation open agentic large language model of the Xing series (formerly TeleChat), developed by China Telecom Artificial Intelligence Technology Co., Ltd. and open-sourced on Hugging Face and ModelScope on September 16, 2026 under Apache 2.0. It is a sparse Mixture-of-Experts model with 29B total parameters and only 4B activated per token — the first model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework — and is deeply optimized for complex engineering (agentic coding) tasks.

Its agent-oriented architecture combines mHC (multi-head hyper-connections) with MLA (Multi-head Latent Attention) and MTP (Multi-Token Prediction): 64 routed experts with 4 active per token plus 1 shared expert (sigmoid scoring, noaux_tc Top-K routing), 2 first-k dense layers, 40 layers, hidden size 3584, dense FFN intermediate size 9216, expert intermediate size 1024, and a 131,072-token vocabulary. MLA compresses the KV cache into a low-rank latent (kv_lora_rank 512, q_lora_rank 768, qk_rope_head_dim 64, qk_nope_head_dim 128, v_head_dim 128); MTP adds one nextn predict layer (num_nextn_predict_layers 1) to accelerate decoding. The mHC hyper-connection mechanism (hc_mult 4, hc_sinkhorn_iters 20) multiplies residual-stream connectivity for stable training and better long-context coherence. A native 256K (262,144-token) context is natively supported and extensible to 512K via YaRN RoPE scaling (factor 64 from original_max_position_embeddings 4096, beta_fast 32 / beta_slow 1).

Training was co-optimized with Ascend NPU (Ascend 910C clusters, MindSpore/MindFormers): fine-grained MoE communication optimization, selective recomputation, DVM automatic graph-operator fusion, and Ascend C mHC fused operators lifted overall training throughput by approximately 96% over out-of-the-box performance. Fine-tuning is supported through LLaMA-Factory and MindFormers; inference through vLLM, SGLang, and KTransformers; agent frameworks such as OpenCode, Claude Code, OpenClaw, and Hermes are explicitly adapted. An OpenAI-compatible API with enable_thinking toggle is documented; recommended sampling is temperature 1.0 / top_p 0.95 for reasoning and 0.8 / 0.95 for coding/agent tasks (repetition_penalty 1.05).

On benchmarks it scores 75.0 on SWE-bench Verified (SWE-agent harness, 210K context), 57.5 on Terminal-Bench 2.1 (terminus-2 harness, avg of 3 runs), 76.55 on Claw-Eval (avg of 3 runs), 64.63 on Tau3-Bench (avg pass^1 over 4 runs), 90.0 on AIME2026 (avg of 5 runs), 69.67 on IFBench, 61.0 on AA-LCR, 66.0 on SWE-bench Multilingual, and 60.8 on DeepresearchBII (OpenCode harness, Exa MCP enabled) — surpassing Gemma4-26B-A4B on most agentic benchmarks at roughly the same active-parameter budget. Weights are distributed in BF16 (~31.2B safetensors parameters, including F32 mHC scaling constants); 4-bit quantization runs locally on an RTX 3090/4090 (~15GB VRAM); community finetunes and quantized GGUF builds exist.

Training Data First model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework (Ascend 910C clusters, MindFormers); deeply optimized for complex engineering tasks with ~96% training throughput improvement via fine-grained MoE communication optimization, selective recomputation, DVM graph-operator fusion, and Ascend C mHC fused operators.

Benchmark Scores

Benchmark Score Date
AA-LCR
long_context
61.00%
16.09.2026
AIME 26
stem_reasoning
90.00%
16.09.2026
TAU3-Bench
general_agent
64.63%
16.09.2026
Claw-Eval Avg
coding_agent
76.55%
16.09.2026
DeepresearchBII
general_agent
60.80%
16.09.2026
SWE-bench Multilingual
coding_agent
66.00%
16.09.2026
IFBench
instruction_following
69.67
16.09.2026
SWE-bench Verified
coding_agent
75.00%
16.09.2026
Terminal-Bench 2.1 (Terminus-2)
coding_agent
57.50%
16.09.2026

Architecture

Decoder Block input Embedding vocab 131K · d 3584 Full Attention MLA · 32 heads ×40 MoE FFN 64 experts · top-4 · +1 shared · dᴻ 1024 MTP Head ×1 speculative layer Final Norm LM Head vocab 131K output
Attention
Multi-head Latent Attention
MoE
64 experts · top-4 per token
Layers
40
Hidden size
3584
Context
262K tokens
RoPE θ
10K
Parameters
31215M
Active params
4000M

Source: Hugging Face config.json · Xing4_0ForCausalLM · model repo

Type: decoder-only transformer
Attention: mla
Decoder: sparse MoE
MoE: yes (64 experts)
Routing: top-k routing with sigmoid scoring (noaux_tc), norm_topk_prob
Layers 40
Context length 262K
Extended context 524K
Attention heads 32
KV heads 32
Hidden size 3584
Vocabulary 131K
FFN dim 9216
MTP layers 1
RoPE θ 10K
Expert Count 64
First K Dense Layers 2
Hc Mult 4
Hc Sinkhorn Iters 20
Layer Pattern first 2 dense layers, then sparse MoE layers
Mla Dims {'kv_lora_rank': 512, 'q_lora_rank': 768, 'qk_nope_head_dim': 128, 'qk_rope_head_dim': 64, 'v_head_dim': 128}
Moe {'enabled': True, 'experts_active': 4, 'experts_total': 64, 'first_k_dense': 2, 'intermediate_size': 1024, 'layer_freq': 1, 'norm_topk_prob': True, 'route_top_k': 4, 'routed_scaling_factor': 2, 'routing_method': 'noaux_tc (auxiliary-loss-free, token-choice)', 'scoring_func': 'sigmoid', 'shared_experts': 1}
Rope Scaling {'beta_fast': 32, 'beta_slow': 1, 'factor': 64, 'mscale': 1, 'mscale_all_dim': 1, 'original_max_position_embeddings': 4096, 'type': 'yarn'}
Tied embeddings No

Related Models