GLM-5.3-Flash

Zhipu AI

Parameters

320.0B total / 18.0B active

MoE: total / active

Architecture

Hybrid sparse + linear attention MoE with mHC

Released

27.08.2026

License

MIT License

Open Weights Commercial Use Multimodal BF16, F8_E4M3, F32 GLM en zh

Input Modalities

text image

Output Modalities

text

Context (native)

1,000,000 tokens

Context (extended)

1,000,000 tokens

Openness Index Score 100.0/100

About

GLM-5.3-Flash (zai-org/GLM-5.3-Flash) is the first natively multimodal model in the GLM-5 series - 320B total parameters with just 18B active, released August 27, 2026 under the MIT License. It outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM-5.3-Flash starts from a newly trained base model with architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series it introduces a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities, and adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with its latest 30T-token multimodal pre-training corpus, these changes let GLM-5.3-Flash deliver more intelligence with less compute. Context length 1,000,000 tokens; text and image input. Blog: z.ai/blog/glm-5.3-flash; GLM-5 technical report: arXiv 2602.15763.

Training Data 30T-token multimodal pre-training corpus

Benchmark Scores

Benchmark Score Date
Terminal Bench 2.1
coding_agent
93.02%
27.08.2026
DeepSWE 1.1
coding_agent
83.69%
27.08.2026
Agents' Last Exam
general_agent
74.06%
27.08.2026
Automation-Bench
general_agent
86.36%
27.08.2026
HLE (with tools)
stem_reasoning
80.51%
27.08.2026
GDPVal-AA v2
general_agent
96.83%
27.08.2026
MMLU
knowledge
93.44%
27.08.2026
BIG-Bench Hard
reasoning
100.00%
27.08.2026
Artificial Analysis Intelligence Index
composite
100.00%
27.08.2026
HellaSwag
reasoning
100.00%
27.08.2026

Serving on Chinese AI Chips

GLM-5.3-Flash was served at scale on Chinese AI chips with a high-bandwidth interconnect and optimized serving stack. Key optimizations include:

  • Intra-node tensor parallelism for Linear Attention and LM head
  • ReplaySSM
  • W8A8 quantization
  • Hybrid INT8/FP8/BF16 cache quantization
  • Layer Split
  • Encode-Prefill-Decode (EPD) disaggregated architecture

A 3x improvement in end-to-end serving performance was achieved compared to baseline, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. Notably, a GLM-5.3-powered infrastructure agent assisted engineers in optimizing the serving stack.

30T-Token Multimodal Pre-training

GLM-5.3-Flash was pre-trained on a 30T-token multimodal corpus, the latest in the GLM series. The combination of optimized pre-training corpus and architectural improvements enables the model to produce more intelligence with less compute.

Cost Efficiency

GLM-5.3-Flash delivers frontier-level intelligence at approximately 1/10th the cost of comparable models. On the Artificial Analysis Intelligence Index v4.1.1, it scores 57 at just $0.045 per task (discounted).

The model was tested anonymously as 'ox-alpha' on OpenCode and OpenRouter before release, becoming the most popular model of the week \u2014 with all traffic served on Chinese AI chips.

Reasoning Effort Control

GLM-5.3-Flash supports controlling the thinking budget through the reasoning_effort parameter, which accepts three levels: low, high, and max. It defaults to max if not passed. For benchmark and leaderboard reproduction, keep the default max.

In the chat template, clear_thinking defaults to false. For chat scenarios, explicitly pass clear_thinking=true.

Natively Multimodal

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. It supports text and image inputs, with multimodal pre-training on a 30T-token corpus.

Hybrid Sparse + Linear Attention Architecture

GLM-5.3-Flash is the first model in the GLM series to introduce a hybrid architecture combining sparse and linear attention. This sharply reduces long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to improve scaling efficiency.\n\nKey architecture stats:\n- Total parameters: 320B\n- Active parameters: 18B (per token)\n- Layers: 45 (compared to 92 in GLM-4.5)\n- Context length: 1M tokens\n\nCompared to GLM-4.5 (355B total / 32B active / 92 layers), GLM-5.3-Flash nearly halves both activated parameters and layer count while delivering superior performance.

Architecture

Decoder Block ×45 input Embedding vocab 155K · d 4096 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×3 Full Attention 64 heads ×1 Linear / Recurrent 64 heads ×1 MoE FFN 288 experts · top-8 · +1 shared · dᴻ 2048 MTP Head ×1 speculative layer Final Norm LM Head vocab 155K output
Attention
Sparse Attention
MoE
288 experts · top-8 per token
Layers
45
Hidden size
4096
Context
1M tokens
Parameters
320000M
Active params
18000M

Source: Hugging Face config.json · Glm5NextForConditionalGeneration · exact layer pattern · model repo

Type: Mixture-of-Experts (MoE) with hybrid sparse + linear attention
Attention: Hybrid sparse + linear attention
Decoder: Transformer
MoE: yes (? experts)
Routing: top-k
Layers 45
Total parameters 320B
Active parameters 18B
Context length 1M
Architecture

Hybrid sparse + linear attention with Manifold-Constrained Hyper-Connections (mHC)

Comparison

GLM-4.5: 355B total / 32B active / 92 layers; GLM-5.3-Flash: 320B total / 18B active / 45 layers

Key Innovation

First GLM model with hybrid sparse + linear attention, reducing long-context serving costs

Training Pipeline

  1. 1
    pretraining

    Multimodal Pre-training

    30T-token multimodal pre-training corpus. First GLM model trained with hybrid sparse + linear attention architecture and Manifold-Constrained Hyper-Connections (mHC).

  2. 2
    other

    Anonymous testing as ox-alpha

    Before release, GLM-5.3-Flash was tested anonymously as 'ox-alpha' on OpenCode and OpenRouter to gather user feedback. It became the most popular model of the week, with all traffic served on Chinese AI chips.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
30T-token multimodal pre-training corpus pretraining 30T tokens text, image mixed

Trend Analysis

24h Change

+3.8%

Current

1,876

huggingface

followers

+0.4%

huggingface

downloads

+16.4%

huggingface

downloads_all_time

+16.4%

ollama

downloads

+25.5%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 49,700 pulls daily 01.09.2026
huggingface downloads_all_time 441,348 daily 01.09.2026
huggingface followers 20,001 daily 01.09.2026
huggingface likes 1,876 daily 01.09.2026
huggingface downloads 441,348 daily 01.09.2026
ollama downloads 39,600 pulls daily 31.08.2026
huggingface downloads_all_time 379,271 daily 31.08.2026
huggingface followers 19,927 daily 31.08.2026
huggingface likes 1,807 daily 31.08.2026
huggingface downloads 379,271 daily 31.08.2026
ollama downloads 28,500 pulls daily 30.08.2026
huggingface downloads_all_time 346,516 daily 30.08.2026
huggingface followers 19,819 daily 30.08.2026
huggingface likes 1,708 daily 30.08.2026
huggingface downloads 346,516 daily 30.08.2026
huggingface followers 19,739 daily 29.08.2026
huggingface likes 1,617 daily 29.08.2026
huggingface downloads 189,793 daily 29.08.2026
huggingface followers 19,618 daily 28.08.2026
huggingface likes 1,499 daily 28.08.2026
huggingface downloads 34 daily 28.08.2026
huggingface followers 19,458 daily 27.08.2026
huggingface likes 1,313 daily 27.08.2026
huggingface downloads 34 daily 27.08.2026

View full metric history →

Related Models