Open-Weight LLM Models — Live Ranking, Benchmarks & Specs
Explore, filter and compare open-weight language models with live specs, benchmarks and license data
DeepSeek-V4-Flash-Vision-Exp
DeepSeek
MoE with vision encoder and aligner, DFlash attention, DSpark speculative decoding
DeepSeek
MIT License
Openness Score
1.0
Spark-X2.5-4B
SparkLLM (XHToken)
Hybrid attention transformer: 1 full-attention layer combined with 3 sliding-window attention (SWA) layers, natively supporting up to 1M-token context
Spark-X2.5
Apache License 2.0
Openness Score
1.0
Spark-X2.5-1.7B
SparkLLM (XHToken)
Hybrid attention transformer: 1 full-attention layer combined with 3 sliding-window attention (SWA) layers, natively supporting up to 1M-token context
Spark-X2.5
Apache License 2.0
Openness Score
1.0
Hy4 preview
Tencent
MoE (78 layers) with Gated DeepSeek Sparse Attention (Gated DSA) + IndexCache cross-layer sparse index reuse, iHC (identity Hyper-Connections) residual streams, 1 native MTP layer (10B total / 0.7B active) for speculative decoding
Hunyuan
Apache License 2.0
Openness Score
1.0
GLM-5.3-Flash
Zhipu AI
Hybrid sparse + linear attention MoE with mHC
GLM
MIT License
Openness Score
1.0
Qwen3.8-Flash-Next
Qwen Team (Alibaba)
Hybrid Attention (Gated DeltaNet + Qwen Sparse Attention) with MoE, N-gram Embedding, and Gated Residual
Qwen
Qwen Community License 1.0
Openness Score
0.7
GLM-5.3
Zhipu AI
MoE with IndexShare sparse attention (DSA) + MTP layer
GLM
GLM-5.3 License
Openness Score
0.7
Ornith-1.5-35B-A3B
Ornith AI
Mixture-of-Experts (qwen3_5_moe)
Ornith-1.5
MIT License
Openness Score
1.0
DeepSeek-V4-Pro-0813
DeepSeek
MoE with DSpark speculative decoding module
DeepSeek
MIT License
Openness Score
1.0
Muse-Glimmer-30B
Meta Superintelligence Lab
Dense Causal Transformer with Perception Encoder
Muse Glimmer
Apache License 2.0
Openness Score
1.0
Qwen3.8-2.4T-A95B
Qwen
92-layer hybrid MoE: 23 x (3 x (Gated DeltaNet -> MoE) + 1 x (Gated Attention -> MoE)); 512 experts (10 routed + 1 shared, dim 2048); hidden 8192; GDN 128V/16QK dim 128; GA 64Q/4KV dim 256; MTP multi-steps; ctx 262144 -> 1010000.
Qwen3.8
Qwen3.8-Max License
Openness Score
0.7
Qwen3.8-27B
Qwen Team (Alibaba Cloud)
Hybrid Gated DeltaNet + Gated Attention (dense, vision-language)
Qwen3.8
Apache License 2.0
Openness Score
1.0
DeepSeek-V4-Flash-0731
DeepSeek
MoE with DSpark speculative decoding module
DeepSeek
MIT License
Openness Score
1.0
Hy3
Tencent
MoE (80 layers + 1 MTP) with GQA (64 heads/8 KV) + QK-Norm, 192 experts top-8 + 1 shared expert, sigmoid router with expert bias
Hunyuan
Apache License 2.0
Openness Score
1.0
Ornith-1.0-35B-A3B
Ornith AI
Mixture-of-Experts (qwen3_5_moe)
Ornith-1.0
MIT License
Openness Score
1.0