Qwen3.5-4B

Qwen

Parameters

4.0B

Architecture

Causal Language Model with Vision Encoder - Hybrid Gated Delta Networks + Gated Attention (dense FFN)

Released

27.02.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16 Qwen3.5 en zh 201 languages

Input Modalities

text image video

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

1,010,000 tokens

Openness Index Score 100.0/100

About

Qwen3.5-4B (Qwen/Qwen3.5-4B) is the small dense member of Alibaba's Qwen3.5 generation - a unified vision-language foundation model with 4B parameters handling text, image and video input in 201 languages and dialects, trained with Early fusion (multimodal pre-training) on multimodal tokens and post-trained with RL scaled across million-agent RL environments.

Its 32-layer hybrid layout repeats 8x (3x (Gated DeltaNet -> FFN) + 1x (Gated Attention -> FFN)): 24 linear-attention layers (Gated DeltaNet with 32 V / 16 QK heads, head dim 128) and 8 full-attention layers (Gated Attention with 16 Q / 4 KV heads, head dim 256, RoPE dim 64), with hidden size 2560, dense FFN 9216 (no MoE), 248K vocabulary tied to token embeddings, and Multi-Token Prediction (MTP) trained with multi-steps. Context is 262,144 tokens natively, extensible up to 1,010,000 tokens.

As a Qwen3.5 family member it delivers the generation's agentic capabilities (native function calling, thinking-mode control via enable_thinking) in a single-GPU-sized package, outperforming much larger Qwen3-VL models across reasoning, coding, agents and visual understanding. Released February 27, 2026 under Apache 2.0.

Training Data Pre-trained on multimodal tokens with early fusion. Post-trained with RL across million-agent environments. 201 languages and dialects.

Benchmark Scores

Benchmark Score Date
VideoMMMU
video_understanding
74.10
01.02.2026
MLVU
video_understanding
68.97%
01.02.2026
MVBench
video_understanding
38.46%
01.02.2026
LVBench
video_understanding
25.00%
01.02.2026
MMVU
video_understanding
64.90
01.02.2026
ScreenSpot Pro
general_agent
60.40%
01.02.2026
OSWorld-Verified
general_agent
35.60
01.02.2026
AndroidWorld
general_agent
58.60
01.02.2026
TIR-Bench
vision_language
40.62%
01.02.2026
V-Star
vision_language
64.57%
01.02.2026
SLAKE
vision_language
79.55%
01.02.2026
PMC-VQA
vision_language
71.11%
01.02.2026
MedXpertQA-MM
vision_language
44.29%
01.02.2026
ParseBench Mean
document_understanding
35.40
01.02.2026
ParseBench Text Content
document_understanding
88.90
01.02.2026
ParseBench Text Formatting
document_understanding
57.80
01.02.2026
Ifstruct V1
instruction_following
36.25
01.02.2026
MMBench EN-DEV v1.1
vision_language
43.33%
01.02.2026
MMLU-ProX
multilingual
3.87%
01.02.2026
SimpleVQA
vision_language
43.40
01.02.2026
NOVA-63
multilingual
35.82%
01.02.2026
HallusionBench
vision_language
50.50%
01.02.2026
INCLUDE
multilingual
71.00
01.02.2026
OmniDocBench 1.5
document_understanding
94.09%
01.02.2026
Global PIQA
multilingual
78.90
01.02.2026
CharXiv (RQ)
document_understanding
18.97%
01.02.2026
PolyMATH
multilingual
25.28%
01.02.2026
MMLongBench-Doc
document_understanding
39.39%
01.02.2026
WMT24++ (en→xx)
multilingual
59.68%
01.02.2026
CC-OCR
document_understanding
61.87%
01.02.2026
MAXIFE
multilingual
78.00
01.02.2026
MMMU
vision_language
83.84%
01.02.2026
OCRBench
document_understanding
54.19%
01.02.2026
MMMU-Pro
vision_language
56.38%
01.02.2026
ERQA
spatial_intelligence
41.90%
01.02.2026
MathVision
vision_language
55.21%
01.02.2026
CountBench
spatial_intelligence
80.77%
01.02.2026
MathVista (mini)
vision_language
68.97%
01.02.2026
RefCOCO (avg)
spatial_intelligence
88.10
01.02.2026
We-Math
vision_language
100.00%
01.02.2026
EmbSpatialBench
spatial_intelligence
74.22%
01.02.2026
RefSpatialBench
spatial_intelligence
77.29%
01.02.2026
ZEROBench
vision_language
3.00
01.02.2026
LingoQA
spatial_intelligence
89.02%
01.02.2026
ZEROBench_sub
vision_language
2.94%
01.02.2026
Hypersim
spatial_intelligence
71.43%
01.02.2026
VlmsAreBlind
vision_language
79.25%
01.02.2026
Nuscene
spatial_intelligence
9.90
01.02.2026
BabyVision
vision_language
4.42%
01.02.2026
VideoMME (w sub.)
video_understanding
47.15%
01.02.2026
RealWorldQA
vision_language
79.07%
01.02.2026
VideoMME (w/o sub.)
video_understanding
18.60%
01.02.2026
MMStar
vision_language
48.11%
01.02.2026
MMLU-Pro
knowledge
61.61%
01.02.2026
MMLU-Redux
knowledge
71.43%
01.02.2026
SuperGPQA
knowledge
60.14%
01.02.2026
GPQA Diamond
stem_reasoning
65.78%
01.02.2026
IFEval
instruction_following
91.36%
01.02.2026
IFBench
instruction_following
60.73%
01.02.2026
Multi-Challenge
instruction_following
30.09%
01.02.2026
AA-LCR
long_context
71.25%
01.02.2026
LongBench v2
long_context
56.79%
01.02.2026
HMMT Feb 25
stem_reasoning
53.78%
01.02.2026
HMMT Nov 25
stem_reasoning
76.80
01.02.2026
LiveCodeBench v6
stem_reasoning
48.57%
01.02.2026
OJBench
stem_reasoning
31.57%
01.02.2026
BFCL-V4
general_agent
47.86%
01.02.2026
TAU2-Bench
general_agent
78.77%
01.02.2026
VITA-Bench
general_agent
45.60%
01.02.2026
DeepPlanning
general_agent
14.43%
01.02.2026
MMMLU
multilingual
37.99%
01.02.2026
TAU3-Bench
general_agent
8.12%
—
MCP-Atlas
general_agent
38.58%
—
Workspace Bench
57.87%
—
BrowseComp
general_agent
12.71%
—
SWE-bench Pro
coding_agent
36.75%
—
SWE-bench Verified
coding_agent
44.04%
—
SWE-bench Multilingual
coding_agent
26.86%
—
SciCode
53.25%
—
Gaokao 2026
90.32%
—
AIME 26
stem_reasoning
79.34%
—
HMMT Feb 26
stem_reasoning
63.73%
—
IMOAnswerBench
stem_reasoning
65.20%
—
Humanity's Last Exam
stem_reasoning
12.31%
—
GPQA
reasoning
83.92%
—
Multi-IF
instruction_following
88.80%
—
Terminal Bench 2.1
coding_agent
28.16%
—
BrowseComp-zh
general_agent
54.50%
—
Gaia2
general_agent
83.76%
—
NoLiMa
long_context
63.61%
—
GDPVal-AA v2
general_agent
0.64%
—
LongBenchPro
long_context
100.00%
—
Claw-Eval Avg
coding_agent
61.36%
—
WildClawBench
coding_agent
18.38%
—
LCB-Pro 25Q2 (Easy)
stem_reasoning
83.19%
—
LCB-Pro 25Q2 (Medium)
stem_reasoning
40.00%
—
AIME 2025
stem_reasoning
78.22%
—
MATH-500
math
100.00%
—
C-Eval
knowledge
26.61%
01.02.2026
AI2D_TEST
document_understanding
87.71%
01.02.2026
MCPMark
general_agent
11.06%
—
DynaMath
vision_language
50.56%
01.02.2026
QwenClawBench
coding_agent
66.67%
—

Model Tree, Spaces and Collection

Model tree for Qwen/Qwen3.5-4B

Base model

Qwen/Qwen3.5-4B-Base

Finetuned

(144)

this model

Adapters

560 models

Finetunes

628 models

Merges

10 models

Quantizations

414 models

Spaces using Qwen/Qwen3.5-4B 100

Collection including Qwen/Qwen3.5-4B

[

Qwen3.5

Collection

21 items • Updated Mar 9 • 1.77k

](https://huggingface.co/collections/Qwen/qwen35)

Citation

Citation

If you find our work helpful, feel free to give us a cite.

@misc{qwen3.5,
    title  = {{Qwen3.5}: Towards Native Multimodal Agents},
    author = {{Qwen Team}},
    month  = {February},
    year   = {2026},
    url    = {https://qwen.ai/blog?id=qwen3.5}
}

Safetensors

Model size

5B params

Tensor type

BF16

·

F32

·

Best Practices

Best Practices

To achieve optimal performance, we recommend the following settings:

  1. Sampling Parameters:

    • We suggest using the following sets of sampling parameters depending on the mode and task type:
      • Thinking mode for general tasks:
        temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
      • Thinking mode for precise coding tasks (e.g., WebDev):
        temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
      • Instruct (or non-thinking) mode for general tasks:
        temperature=0.7, top_p=0.8, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
      • Instruct (or non-thinking) mode for reasoning tasks:
        temperature=1.0, top_p=1.0, top_k=40, min_p=0.0, presence_penalty=2.0, repetition_penalty=1.0
    • For supported frameworks, you can adjust the presence_penalty parameter between 0 and 2 to reduce endless repetitions. However, using a higher value may occasionally result in language mixing and a slight decrease in model performance.
  2. Adequate Output Length: We recommend using an output length of 32,768 tokens for most queries. For benchmarking on highly complex problems, such as those found in math and programming competitions, we suggest setting the max output length to 81,920 tokens. This provides the model with sufficient space to generate detailed and comprehensive responses, thereby enhancing its overall performance.

  3. Standardize Output Format: We recommend using prompts to standardize model outputs when benchmarking.

    • Math Problems: Include "Please reason step by step, and put your final answer within \boxed{}." in the prompt.
    • Multiple-Choice Questions: Add the following JSON structure to the prompt to standardize responses: "Please show your choice in the answer field with only the choice letter, e.g., "answer": "C"."
  4. No Thinking Content in History: In multi-turn conversations, the historical model output should only include the final output part and does not need to include the thinking content. It is implemented in the provided chat template in Jinja2. However, for frameworks that do not directly use the Jinja2 chat template, it is up to the developers to ensure that the best practice is followed.

  5. Long Video Understanding: To optimize inference efficiency for plain text and images, the size parameter in the released video_preprocessor_config.json is conservatively configured. It is recommended to set the longest_edge parameter in the video_preprocessor_config file to 469,762,048 (corresponding to 224k video tokens) to enable higher frame-rate sampling for hour-scale videos and thereby achieve superior performance. For example,

    {"longest_edge": 469762048, "shortest_edge": 4096}
    
    

    Alternatively, override the default values via engine startup parameters. For implementation details, refer to: vLLM / SGLang.

Processing Ultra-Long Texts (1M context)

Processing Ultra-Long Texts

Qwen3.5 natively supports context lengths of up to 262,144 tokens. For long-horizon tasks where the total length (including both input and output) exceeds this limit, we recommend using RoPE scaling techniques to handle long texts effectively., e.g., YaRN.

YaRN is currently supported by several inference frameworks, e.g., transformers, vllm, ktransformers and sglang. In general, there are two approaches to enabling YaRN for supported frameworks:

  • Modifying the model configuration file: In the config.json file, change the rope_parameters fields in text_config to:

    {
        "mrope_interleaved": true,
        "mrope_section": [
            11,
            11,
            10
        ],
        "rope_type": "yarn",
        "rope_theta": 10000000,
        "partial_rotary_factor": 0.25,
        "factor": 4.0,
        "original_max_position_embeddings": 262144,
    }
    
    
  • Passing command line arguments:

    For vllm, you can use

    VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1010000  
    
    

    For sglang and ktransformers, you can use

    SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 python -m sglang.launch_server ... --json-model-override-args '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --context-length 1010000
    
    

All the notable open-source frameworks implement static YaRN, which means the scaling factor remains constant regardless of input length, potentially impacting performance on shorter texts. We advise modifying the rope_parameters configuration only when processing long contexts is required. It is also recommended to modify the factor as needed. For example, if the typical context length for your application is 524,288 tokens, it would be better to set factor as 2.0.

Agentic Usage (Qwen-Agent, Qwen Code)

Agentic Usage

Qwen3.5 excels in tool calling capabilities.

Qwen-Agent

We recommend using Qwen-Agent to quickly build Agent applications with Qwen3.5.

To define the available tools, you can use the MCP configuration file, use the integrated tool of Qwen-Agent, or integrate other tools by yourself.

import os
from qwen_agent.agents import Assistant

## Define LLM
## Using Alibaba Cloud Model Studio
llm_cfg = {
    # Use the OpenAI-compatible model service provided by DashScope:
    'model': 'Qwen3.5-4B',
    'model_type': 'qwenvl_oai',
    'model_server': 'https://dashscope.aliyuncs.com/compatible-mode/v1',
    'api_key': os.getenv('DASHSCOPE_API_KEY'),

    'generate_cfg': {
        'use_raw_api': True,
        # When using Dash Scope OAI API, pass the parameter of whether to enable thinking mode in this way
        'extra_body': {
            'enable_thinking': True
        },
    },
}

## Using OpenAI-compatible API endpoint.
## functionality of the deployment frameworks and let Qwen-Agent automate the related operations.
##
## llm_cfg = {
##     # Use your own model service compatible with OpenAI API by vLLM/SGLang:
##     'model': 'Qwen/Qwen3.5-4B',
##     'model_type': 'qwenvl_oai',
##     'model_server': 'http://localhost:8000/v1',  # api_base
##     'api_key': 'EMPTY',
##
##     'generate_cfg': {
##         'use_raw_api': True,
##         # When using vLLM/SGLang OAI API, pass the parameter of whether to enable thinking mode in this way
##         'extra_body': {
##             'chat_template_kwargs': {'enable_thinking': True}
##         },
##     },
## }

## Define Tools
tools = [
    {'mcpServers': {  # You can specify the MCP configuration file
            "filesystem": {
                "command": "npx",
                "args": ["-y", "@modelcontextprotocol/server-filesystem", "/Users/xxxx/Desktop"]
            }
        }
    }
]

## Define Agent
bot = Assistant(llm=llm_cfg, function_list=tools)

## Streaming generation
messages = [{'role': 'user', 'content': 'Help me organize my desktop.'}]
for responses in bot.run(messages=messages):
    pass
print(responses)

## Streaming generation
messages = [{'role': 'user', 'content': 'Develop a dog website and save it on the desktop'}]
for responses in bot.run(messages=messages):
    pass
print(responses)

Qwen Code

Qwen Code is an open-source AI agent for the terminal, optimized for Qwen models. It helps you understand large codebases, automate tedious work, and ship faster.

For more information, please refer to Qwen Code.

Serving Qwen3.5 (vLLM/SGLang, chat completions)

Serving Qwen3.5

Qwen3.5 can be served via APIs with popular inference frameworks. In the following, we show example commands to launch OpenAI-Compatible API servers for Qwen3.5 models.

Inference efficiency and throughput vary significantly across frameworks. We recommend using the latest framework versions to ensure optimal performance and compatibility. For production workloads or high-throughput scenarios, dedicated serving engines such as SGLang, KTransformers or vLLM are strongly recommended.

The model has a default context length of 262,144 tokens. If you encounter out-of-memory (OOM) errors, consider reducing the context window. However, because Qwen3.5 leverages extended context for complex tasks, we advise maintaining a context length of at least 128K tokens to preserve thinking capabilities.

SGLang

SGLang is a fast serving framework for large language models and vision language models. SGLang from the main branch of the open-source repository is required for Qwen3.5, which can be installed using the following command in a fresh environment:

uv pip install 'git+https://github.com/sgl-project/sglang.git#subdirectory=python&egg=sglang[all]'

See its documentation for more details.

The following will create API endpoints at http://localhost:8000/v1:

  • Standard Version: The following command can be used to create an API endpoint with maximum context length 262,144 tokens using tensor parallel on 8 GPUs.

    python -m sglang.launch_server --model-path Qwen/Qwen3.5-4B --port 8000 --tp-size 1 --mem-fraction-static 0.8 --context-length 262144 --reasoning-parser qwen3
    
    
  • Tool Use: To support tool use, you can use the following command.

    python -m sglang.launch_server --model-path Qwen/Qwen3.5-4B --port 8000 --tp-size 1 --mem-fraction-static 0.8 --context-length 262144 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
    
    
  • Multi-Token Prediction (MTP): The following command is recommended for MTP:

    python -m sglang.launch_server --model-path Qwen/Qwen3.5-4B --port 8000 --tp-size 1 --mem-fraction-static 0.8 --context-length 262144 --reasoning-parser qwen3 --speculative-algo NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
    
    

vLLM

vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. vLLM from the main branch of the open-source repository is required for Qwen3.5, which can be installed using the following command in a fresh environment:

uv pip install vllm --torch-backend=auto --extra-index-url https://wheels.vllm.ai/nightly

See its documentation for more details.

For detailed Qwen3.5 usage guide, see the vLLM Qwen3.5 recipe.

The following will create API endpoints at http://localhost:8000/v1:

  • Standard Version: The following command can be used to create an API endpoint with maximum context length 262,144 tokens using tensor parallel on 8 GPUs.

    vllm serve Qwen/Qwen3.5-4B --port 8000 --tensor-parallel-size 1 --max-model-len 262144 --reasoning-parser qwen3 
    
    
  • Tool Call: To support tool use, you can use the following command.

    vllm serve Qwen/Qwen3.5-4B --port 8000 --tensor-parallel-size 1 --max-model-len 262144 --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder 
    
    
  • Multi-Token Prediction (MTP): The following command is recommended for MTP:

    vllm serve Qwen/Qwen3.5-4B --port 8000 --tensor-parallel-size 1 --max-model-len 262144 --reasoning-parser qwen3 --speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'
    
    
  • Text-Only: The following command skips the vision encoder and multimodal profiling to free up memory for additional KV cache:

    vllm serve Qwen/Qwen3.5-4B --port 8000 --tensor-parallel-size 1 --max-model-len 262144 --reasoning-parser qwen3 --language-model-only
    
    

KTransformers

KTransformers is a flexible framework for experiencing cutting-edge LLM inference optimizations with CPU-GPU heterogeneous computing. For running Qwen3.5 with KTransformers, see the KTransformers Deployment Guide.

Hugging Face Transformers

Hugging Face Transformers contains a lightweight server which can be used for quick testing and moderate load deployment. The latest transformers is required for Qwen3.5:

pip install "transformers[serving] @ git+https://github.com/huggingface/transformers.git@main"

See its documentation for more details. Please also make sure torchvision and pillow are installed.

Then, run transformers serve to launch a server with API endpoints at http://localhost:8000/v1; it will place the model on accelerators if available:

transformers serve --force-model Qwen/Qwen3.5-4B --port 8000 --continuous-batching

Using Qwen3.5 via the Chat Completions API

The chat completions API is accessible via standard HTTP requests or OpenAI SDKs. Here, we show examples using the OpenAI Python SDK.

Before starting, make sure it is installed and the API key and the API base URL is configured, e.g.:

pip install -U openai

## Set the following accordingly
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"

We recommend using the following set of sampling parameters for generation

  • Thinking mode for general tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
  • Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
  • Instruct (or non-thinking) mode for general tasks: temperature=0.7, top_p=0.8, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
  • Instruct (or non-thinking) mode for reasoning tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0

Please note that the support for sampling parameters varies according to inference frameworks.

Text-Only Input

from openai import OpenAI
## Configured by environment variables
client = OpenAI()

messages = [
    {"role": "user", "content": "Type \"I love Qwen3.5\" backwards"},
]

chat_response = client.chat.completions.create(
    model="Qwen/Qwen3.5-4B",
    messages=messages,
    max_tokens=81920,
    temperature=1.0,
    top_p=0.95,
    presence_penalty=1.5,
    extra_body={
        "top_k": 20,
    }, 
)
print("Chat response:", chat_response)

Image Input

from openai import OpenAI
## Configured by environment variables
client = OpenAI()

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image_url",
                "image_url": {
                    "url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/CI_Demo/mathv-1327.jpg"
                }
            },
            {
    

(table continues in the model card)

Quickstart

Quickstart

Qwen3.5 models operate in thinking mode by default, generating thinking content signified by <think>\n...</think>\n\n before producing the final responses. To disable thinking content and obtain direct response, refer to the examples here.

For streamlined integration, we recommend using Qwen3.5 via APIs. Below is a guide to use Qwen3.5 via OpenAI-compatible API.

Vision-Language Benchmark Results

Vision Language

GPT-5-Nano-2025-08-07 Gemini-2.5-Flash-Lite Qwen3-VL-30B-A3B Qwen3.5-9B Qwen3.5-4B
STEM and Puzzle
MMMU 75.8 73.4 76.0 78.4 77.6
MMMU-Pro 57.2 59.7 63.0 70.1 66.3
MathVision 62.2 52.1 65.7 78.9 74.6
Mathvista(mini) 71.5 72.8 81.9 85.7 85.1
We-Math 62.5 32.1 70.0 75.2 75.4
DynaMath 78.0 69.9 80.1 83.6 83.3
ZEROBench 1.0 1.0 0.0 3.0 3.0
ZEROBench_sub 22.2 19.2 23.7 31.1 26.3
VlmsAreBlind 66.7 68.4 72.5 93.7 92.6
BabyVision 14.4 17.5 18.6 28.6/25.8 16.0/19.1
General VQA
RealWorldQA 71.8 72.2 77.4 80.3 79.5
MMStar 68.6 69.1 75.5 79.7 78.3
MMBenchEN-DEV-v1.1 80.3 82.7 88.9 90.1 89.4
SimpleVQA 46.0 54.1 54.3 51.2 43.4
HallusionBench 58.4 64.5 66.0 69.3 65.0
Text Recognition and Document Understanding
OmniDocBench1.5 55.9 79.4 86.8 87.7 86.2
CharXiv(RQ) 50.1 56.1 56.6 73.0 70.8
MMLongBench-Doc 31.8 46.5 47.4 57.7 54.2
CC-OCR 58.9 72.9 77.8 79.3 76.7
AI2D_TEST 81.9 85.7 86.9 90.2 89.6
OCRBench 75.3 82.5 83.9 89.2 85.0
Spatial Intelligence
ERQA 45.8 44.3 45.3 55.5 54.0
CountBench 80.0 79.2 90.0 97.2 96.3
RefCOCO(avg) -- -- 89.3 89.7 88.1
EmbSpatialBench 74.2 66.1 80.6 83.0 81.3
RefSpatialBench 12.6 11.2 54.2 58.5 54.6
LingoQA 57.0 17.8 62.0 80.4 74.4
Hypersim -- -- 11.4 13.5 12.5
Nuscene -- -- 10.3 11.8 9.9
Video Understanding
VideoMME(w sub.) 71.7 74.6 79.9 84.5 83.5
VideoMME(w/o sub.) 66.2 72.7 73.3 78.4 76.9
VideoMMMU 63.0 69.2 75.0 78.9 74.1
MLVU 69.2 78.5 78.9 84.4 82.8
MVBench -- -- 72.0 74.4 71.2
LVBench -- 60.9 59.2 70.0 66.4
MMVU 63.1 65.3 66.1 67.8 64.9
Visual Agent
ScreenSpot Pro -- -- 60.5 65.2 60.3
OSWorld-Verified -- -- 30.6 41.8 35.6
AndroidWorld -- -- 55.0 57.8 58.6
Tool Calling
TIR-Bench 18.5 21.5 22.5 45.6/31.9 38.9/29.9
V* 68.1 69.6 83.2 90.1/88.5 84.3/86.4
Medical VQA
SLAKE 57.0 65.0 68.8 79.0 76.1
PMC-VQA 37.8 48.8 51.5 57.9 55.5
MedXpertQA-MM 26.7 35.3 35.5 49.9 42.9

* MathVision: our model’s score is evaluated using a fixed prompt, e.g., “Please reason step by step, and put your final answer within \boxed{}.” For other models, we report the higher score between runs with and without the \boxed{} formatting.
* BabyVision: scores reported as "with CI / without CI".
* TIR-Bench and V*: scores reported as "with CI / without CI".
* Empty cells (--) indicate scores not yet available or not applicable.

Language Benchmark Results

Language

GPT-OSS-120B GPT-OSS-20B Qwen3-Next-80B-A3B-Thinking Qwen3-30BA3B-Thinking-2507 Qwen3.5-9B Qwen3.5-4B
Knowledge & STEM
MMLU-Pro 80.8 74.8 82.7 80.9 82.5 79.1
MMLU-Redux 91.0 87.8 92.5 91.4 91.1 88.8
C-Eval 76.2 71.4 89.7 87.4 88.2 85.1
SuperGPQA 54.6 48.5 60.8 56.8 58.2 52.9
GPQA Diamond 80.1 71.5 77.2 73.4 81.7 76.2
Instruction Following
IFEval 88.9 88.2 88.9 88.9 91.5 89.8
IFBench 69.0 65.1 61.5 51.5 64.5 59.2
MultiChallenge 45.3 40.1 51.3 46.5 54.5 49.0
Long Context
AA-LCR 50.7 30.7 51.7 49.0 63.0 57.0
LongBench v2 48.2 45.6 48.0 44.8 55.2 50.0
Reasoning & Coding
HMMT Feb 25 90.0 76.7 73.7 63.1 83.2 74.0
HMMT Nov 25 90.0 81.8 81.2 73.8 82.9 76.8
LiveCodeBench v6 82.7 74.6 68.7 66.0 65.6 55.8
OJBench 41.5 36.3 29.7 25.1 29.2 24.1
General Agent
BFCL-V4 -- -- 49.7 42.4 66.1 50.3
TAU2-Bench -- -- 57.4 41.9 79.1 79.9
VITA-Bench -- -- 29.5 14.1 29.8 22.0
DeepPlanning -- -- 0.4 4.9 18.0 17.6
Multilingualism
MMMLU 78.2 69.7 81.3 78.4 81.2 76.1
MMLU-ProX 74.5 67.3 73.6 69.1 76.3 71.5
NOVA-63 51.1 48.7 53.3 52.5 55.9 54.3
INCLUDE 74.0 65.3 78.3 74.4 75.6 71.0
Global PIQA 84.1 79.8 83.5 80.2 83.2 78.9
PolyMATH 54.0 30.9 62.4 52.6 57.3 51.1
WMT24++ 74.4 67.8 57.4 69.3 72.6 66.6
MAXIFE 83.7 80.1 79.9 77.4 83.4 78.0

* TAU2-Bench: we follow the official setup except for the airline domain, where all models are evaluated by applying the fixes proposed in the Claude Opus 4.5 system card.

* MMLU-ProX: we report the averaged accuracy on 29 languages.
* WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL.
* MAXIFE: we report the accuracy on English + multilingual original prompts (totally 23 settings).
* Empty cells (--) indicate scores not yet available or not applicable.

Model Overview (architecture)

Model Overview

  • Type: Causal Language Model with Vision Encoder
  • Training Stage: Pre-training & Post-training
  • Language Model
    • Number of Parameters: 4B
    • Hidden Dimension: 2560
    • Token Embedding: 248320 (Padded)
    • Number of Layers: 32
    • Hidden Layout: 8 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
    • Gated DeltaNet:
      • Number of Linear Attention Heads: 32 for V and 16 for QK
      • Head Dimension: 128
    • Gated Attention:
      • Number of Attention Heads: 16 for Q and 4 for KV
      • Head Dimension: 256
      • Rotary Position Embedding Dimension: 64
    • Feed Forward Network:
      • Intermediate Dimension: 9216
    • LM Output: 248320 (Tied to token embedding)
    • MTP: trained with multi-steps
  • Context Length: 262,144 natively and extensible up to 1,010,000 tokens.

Qwen3.5 Highlights

Qwen3.5 Highlights

Qwen3.5 features the following enhancement:

  • Unified Vision-Language Foundation: Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks.

  • Efficient Hybrid Architecture: Gated Delta Networks combined with sparse Mixture-of-Experts deliver high-throughput inference with minimal latency and cost overhead.

  • Scalable RL Generalization: Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.

  • Global Linguistic Coverage: Expanded support to 201 languages and dialects, enabling inclusive, worldwide deployment with nuanced cultural and regional understanding.

  • Next-Generation Training Infrastructure: Near-100% multimodal training efficiency compared to text-only training and asynchronous RL frameworks supporting massive-scale agent scaffolds and environment orchestration.

Benchmark Results

For more details, please refer to our blog post Qwen3.5.

Architecture

Decoder Block ×32 input Embedding vocab 248K · d 2560 Linear / Recurrent Hybrid 16:4 · dₕ 256 ×24 Full Attention Hybrid 16:4 · dₕ 256 ×8 Dense FFN silu · d 9216 Final Norm LM Head tied with embedding output
Attention
Hybrid Attention (16:4)
Layers
32
Hidden size
2560
Context
262K tokens
Parameters
4000M

Source: Hugging Face config.json · Qwen3_5ForConditionalGeneration · exact layer pattern · model repo

Type: Hybrid Gated Delta Networks + Gated Attention (dense)
Attention: Gated Attention with RoPE (16 Q heads, 4 KV heads, head_dim=256, RoPE dim=64) + Gated DeltaNet linear attention (32 V heads, 16 QK heads, head_dim=128)
Decoder: Causal Language Model
Routing: sparse
Layers 32
Context length 262K
Extended context 1M
Hidden size 2560
FFN dim 9216
Vision Yes
Hidden Layout 8 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
MTP Yes
RoPE dim 64
Lm Output

248K

Lm Output Note

tied to token embedding

Token Embedding

248K

Token Embedding Note

padded

Training Pipeline

  1. 1
    pretraining

    Pre-training with Multimodal Early Fusion

    Pre-training on multimodal tokens with early fusion. Near-100% multimodal training efficiency compared to text-only training. 201 languages and dialects supported.

  2. 2
    rl

    Reinforcement Learning Post-training

    Post-training with reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability. Asynchronous RL frameworks supporting massive-scale agent scaffolds and environment orchestration.

  3. 3
    other

    Post-training (Thinking + Instruct Modes)

    Post-training stage producing the final model weights compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers. Includes thinking mode and instruct mode capabilities.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Early-fusion multimodal token corpus (text + image) pretraining — —
Million-agent RL task environments rl — —

Trend Analysis

24h Change

+0.2%

7d Change

+1.9%

Current

101,994

huggingface

likes

+0.2%

huggingface

downloads

-0.8%

huggingface

downloads_all_time

+0.6%

ollama

downloads

+0.5%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 19,000,000 pulls daily 01.09.2026
huggingface downloads_all_time 41,140,758 daily 01.09.2026
huggingface followers 101,994 daily 01.09.2026
huggingface likes 867 daily 01.09.2026
huggingface downloads 7,443,149 daily 01.09.2026
ollama downloads 18,900,000 pulls daily 31.08.2026
huggingface downloads_all_time 40,883,935 daily 31.08.2026
huggingface followers 101,772 daily 31.08.2026
huggingface likes 865 daily 31.08.2026
huggingface downloads 7,499,603 daily 31.08.2026
ollama downloads 18,800,000 pulls daily 30.08.2026
huggingface downloads_all_time 40,721,622 daily 30.08.2026
huggingface followers 101,528 daily 30.08.2026
huggingface likes 865 daily 30.08.2026
huggingface downloads 7,646,199 daily 30.08.2026
ollama downloads 18,800,000 pulls daily 29.08.2026
huggingface followers 101,306 daily 29.08.2026
huggingface likes 862 daily 29.08.2026
huggingface downloads 7,438,709 daily 29.08.2026
ollama downloads 18,700,000 pulls daily 28.08.2026
huggingface followers 101,117 daily 28.08.2026
huggingface likes 859 daily 28.08.2026
huggingface downloads 7,268,750 daily 28.08.2026
ollama downloads 18,600,000 pulls daily 27.08.2026
huggingface followers 100,862 daily 27.08.2026
huggingface likes 856 daily 27.08.2026
huggingface downloads 7,533,463 daily 27.08.2026
ollama downloads 18,500,000 pulls daily 26.08.2026
huggingface followers 100,537 daily 26.08.2026
huggingface likes 852 daily 26.08.2026
huggingface downloads 7,750,908 daily 26.08.2026
ollama downloads 18,400,000 pulls daily 25.08.2026
huggingface followers 100,133 daily 25.08.2026
huggingface likes 848 daily 25.08.2026
huggingface downloads 7,782,618 daily 25.08.2026
ollama downloads 18,300,000 pulls daily 24.08.2026
huggingface followers 99,896 daily 24.08.2026
huggingface likes 843 daily 24.08.2026
huggingface downloads 7,695,410 daily 24.08.2026
huggingface followers 99,669 daily 23.08.2026
huggingface likes 841 daily 23.08.2026
huggingface downloads 7,670,188 daily 23.08.2026
huggingface followers 99,461 daily 22.08.2026
huggingface likes 840 daily 22.08.2026
huggingface downloads 7,742,677 daily 22.08.2026
huggingface followers 99,274 daily 21.08.2026
huggingface likes 837 daily 21.08.2026
huggingface downloads 7,715,507 daily 21.08.2026
huggingface followers 99,037 daily 20.08.2026
huggingface likes 835 daily 20.08.2026

View full metric history →

Related Models