Parameters
320.0B total / 18.0B active
MoE: total / active
Architecture
Hybrid sparse + linear attention MoE with mHC
Released
27.08.2026
License
MIT License
Input Modalities
Output Modalities
Context (native)
1,000,000 tokens
Context (extended)
1,000,000 tokens
About
GLM-5.3-Flash (zai-org/GLM-5.3-Flash) is the first natively multimodal model in the GLM-5 series - 320B total parameters with just 18B active, released August 27, 2026 under the MIT License. It outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.
GLM-5.3-Flash starts from a newly trained base model with architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series it introduces a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities, and adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with its latest 30T-token multimodal pre-training corpus, these changes let GLM-5.3-Flash deliver more intelligence with less compute. Context length 1,000,000 tokens; text and image input. Blog: z.ai/blog/glm-5.3-flash; GLM-5 technical report: arXiv 2602.15763.
Training Data 30T-token multimodal pre-training corpus
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
Terminal Bench 2.1
coding_agent
|
93.02%
|
27.08.2026 |
|
DeepSWE 1.1
coding_agent
|
83.69%
|
27.08.2026 |
|
Agents' Last Exam
general_agent
|
74.06%
|
27.08.2026 |
|
Automation-Bench
general_agent
|
86.36%
|
27.08.2026 |
|
HLE (with tools)
stem_reasoning
|
80.51%
|
27.08.2026 |
|
GDPVal-AA v2
general_agent
|
96.83%
|
27.08.2026 |
|
MMLU
knowledge
|
93.44%
|
27.08.2026 |
|
BIG-Bench Hard
reasoning
|
100.00%
|
27.08.2026 |
|
Artificial Analysis Intelligence Index
composite
|
100.00%
|
27.08.2026 |
|
HellaSwag
reasoning
|
100.00%
|
27.08.2026 |
Serving on Chinese AI Chips
GLM-5.3-Flash was served at scale on Chinese AI chips with a high-bandwidth interconnect and optimized serving stack. Key optimizations include:
- Intra-node tensor parallelism for Linear Attention and LM head
- ReplaySSM
- W8A8 quantization
- Hybrid INT8/FP8/BF16 cache quantization
- Layer Split
- Encode-Prefill-Decode (EPD) disaggregated architecture
A 3x improvement in end-to-end serving performance was achieved compared to baseline, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. Notably, a GLM-5.3-powered infrastructure agent assisted engineers in optimizing the serving stack.
30T-Token Multimodal Pre-training
GLM-5.3-Flash was pre-trained on a 30T-token multimodal corpus, the latest in the GLM series. The combination of optimized pre-training corpus and architectural improvements enables the model to produce more intelligence with less compute.
Cost Efficiency
GLM-5.3-Flash delivers frontier-level intelligence at approximately 1/10th the cost of comparable models. On the Artificial Analysis Intelligence Index v4.1.1, it scores 57 at just $0.045 per task (discounted).
The model was tested anonymously as 'ox-alpha' on OpenCode and OpenRouter before release, becoming the most popular model of the week \u2014 with all traffic served on Chinese AI chips.
Reasoning Effort Control
GLM-5.3-Flash supports controlling the thinking budget through the reasoning_effort parameter, which accepts three levels: low, high, and max. It defaults to max if not passed. For benchmark and leaderboard reproduction, keep the default max.
In the chat template, clear_thinking defaults to false. For chat scenarios, explicitly pass clear_thinking=true.
Natively Multimodal
GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. It supports text and image inputs, with multimodal pre-training on a 30T-token corpus.
Hybrid Sparse + Linear Attention Architecture
GLM-5.3-Flash is the first model in the GLM series to introduce a hybrid architecture combining sparse and linear attention. This sharply reduces long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to improve scaling efficiency.\n\nKey architecture stats:\n- Total parameters: 320B\n- Active parameters: 18B (per token)\n- Layers: 45 (compared to 92 in GLM-4.5)\n- Context length: 1M tokens\n\nCompared to GLM-4.5 (355B total / 32B active / 92 layers), GLM-5.3-Flash nearly halves both activated parameters and layer count while delivering superior performance.
Architecture
- Attention
- Sparse Attention
- MoE
- 288 experts · top-8 per token
- Layers
- 45
- Hidden size
- 4096
- Context
- 1M tokens
- Parameters
- 320000M
- Active params
- 18000M
Source: Hugging Face config.json · Glm5NextForConditionalGeneration · exact layer pattern · model repo
Hybrid sparse + linear attention with Manifold-Constrained Hyper-Connections (mHC)
GLM-4.5: 355B total / 32B active / 92 layers; GLM-5.3-Flash: 320B total / 18B active / 45 layers
First GLM model with hybrid sparse + linear attention, reducing long-context serving costs
Training Pipeline
-
1
pretraining
Multimodal Pre-training
30T-token multimodal pre-training corpus. First GLM model trained with hybrid sparse + linear attention architecture and Manifold-Constrained Hyper-Connections (mHC).
-
2
other
Anonymous testing as ox-alpha
Before release, GLM-5.3-Flash was tested anonymously as 'ox-alpha' on OpenCode and OpenRouter to gather user feedback. It became the most popular model of the week, with all traffic served on Chinese AI chips.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| 30T-token multimodal pre-training corpus | pretraining | 30T tokens | text, image | mixed |
Linked Resources
GLM-5: from Vibe Coding to Agentic Engineering
https://arxiv.org/abs/2602.15763
GLM-5.3-Flash: Frontier Intelligence, Flash Cost
https://z.ai/blog/glm-5.3-flash
Z.ai API Platform - GLM-5.3-Flash
https://docs.z.ai/guides/llm/glm-5.3-flash
SGLang - GLM-5.3-Flash cookbook
https://cookbook.sglang.io/autoregressive/GLM/GLM-5.3-Flash
vLLM - GLM-5.3-Flash recipes
https://recipes.vllm.ai/zai-org/GLM-5.3-Flash
TokenSpeed - GLM-5.3-Flash
https://lightseek.org/tokenspeed/recipes/models#glm-5-3-flash
Transformers - GLM5-next docs
https://github.com/huggingface/transformers/blob/main/docs/source/en/model_doc/glm5_next.md
KTransformers - GLM-5.3-Flash Tutorial
https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/GLM-5.3-Flash-Tutorial.md
Unsloth - GLM-5.3 guide
https://unsloth.ai/docs/models/glm-5.3
SGLang repository
https://github.com/sgl-project/sglang
vLLM repository
https://github.com/vllm-project/vllm
TokenSpeed repository
https://github.com/lightseekorg/tokenspeed
KTransformers repository
https://github.com/kvcache-ai/ktransformers
Unsloth repository
https://github.com/unslothai/unsloth
Discord community
https://discord.gg/QR7SARHRxK
Z.ai Coding Plan subscription
https://z.ai/subscribe
ZCode - GLM-5.3-Flash multimodal
https://zcode.z.ai
Trend Analysis
24h Change
+3.8%
Current
1,876
followers
+0.4%
downloads
+16.4%
downloads_all_time
+16.4%
downloads
+25.5%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 49,700 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 441,348 | daily | 01.09.2026 |
| huggingface | followers | 20,001 | daily | 01.09.2026 |
| huggingface | likes | 1,876 | daily | 01.09.2026 |
| huggingface | downloads | 441,348 | daily | 01.09.2026 |
| ollama | downloads | 39,600 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 379,271 | daily | 31.08.2026 |
| huggingface | followers | 19,927 | daily | 31.08.2026 |
| huggingface | likes | 1,807 | daily | 31.08.2026 |
| huggingface | downloads | 379,271 | daily | 31.08.2026 |
| ollama | downloads | 28,500 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 346,516 | daily | 30.08.2026 |
| huggingface | followers | 19,819 | daily | 30.08.2026 |
| huggingface | likes | 1,708 | daily | 30.08.2026 |
| huggingface | downloads | 346,516 | daily | 30.08.2026 |
| huggingface | followers | 19,739 | daily | 29.08.2026 |
| huggingface | likes | 1,617 | daily | 29.08.2026 |
| huggingface | downloads | 189,793 | daily | 29.08.2026 |
| huggingface | followers | 19,618 | daily | 28.08.2026 |
| huggingface | likes | 1,499 | daily | 28.08.2026 |
| huggingface | downloads | 34 | daily | 28.08.2026 |
| huggingface | followers | 19,458 | daily | 27.08.2026 |
| huggingface | likes | 1,313 | daily | 27.08.2026 |
| huggingface | downloads | 34 | daily | 27.08.2026 |