Parameters
35.0B total / 3.0B active
MoE: total / active
Architecture
Hybrid Gated DeltaNet + Gated Attention MoE - Language World Model (qwen3_5_moe)
Released
22.06.2026
License
Apache License 2.0
Input Modalities
Output Modalities
Context (native)
262,144 tokens
Context (extended)
262,144 tokens
About
Qwen-AgentWorld-35B-A3B (Qwen/Qwen-AgentWorld-35B-A3B) is Alibaba's open Language World Model (LWM) - an agent foundation model built as a native world model rather than a post-hoc-adapted general LLM. A single model covers seven unified domains - MCP (tool calling), Search, Terminal, SWE, Android, Web and OS - spanning both text and GUI interaction environments. It generalizes zero-shot to out-of-distribution environments (e.g. OpenClaw), supports controllable perturbations, and can construct fictional training worlds that surpass real-environment training.
Built on Qwen3.5-35B-A3B-Base (35B total, 3B activated; 40 layers in a 10x (3x (Gated DeltaNet -> MoE) + 1x (Gated Attention -> MoE)) hybrid layout with 256 experts, 8 routed + 1 shared per token, hidden 2048), it is trained with a three-stage pipeline: Continual Pre-Training for environment knowledge injection, SFT with next-state-prediction reasoning, and RL with GSPO for simulation fidelity. LWM RL warm-up on single-turn, non-agentic trajectories transfers to multi-turn, tool-calling agentic tasks across 7 benchmarks including 3 entirely out-of-domain. A key disclaimer: no outputs from external API services are included in the training pipeline. Context length is 262,144 tokens; released June 22, 2026 under Apache 2.0. Paper: "Qwen-AgentWorld: Language World Models for General Agents" (arXiv 2606.24597).
Training Data Three-stage pipeline: CPT (environment knowledge injection) → SFT (next-state-prediction reasoning) → RL (GSPO, simulation fidelity). No external API service outputs included in training. Base model: Qwen3.5-35B-A3B-Base.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
AgentWorldBench MCP
general_agent
|
65.00%
|
24.06.2026 |
|
AgentWorldBench Search
general_agent
|
92.88%
|
24.06.2026 |
|
AgentWorldBench Terminal
general_agent
|
70.27%
|
24.06.2026 |
|
AgentWorldBench SWE
general_agent
|
90.79%
|
24.06.2026 |
|
AgentWorldBench Web
general_agent
|
32.41%
|
24.06.2026 |
|
AgentWorldBench OS
general_agent
|
69.27%
|
24.06.2026 |
|
AgentWorldBench Overall
general_agent
|
81.57%
|
24.06.2026 |
|
AgentWorldBench Android
general_agent
|
61.78%
|
24.06.2026 |
Thinking Mode
The model uses thinking mode by default (<think>...</think>) to reason about environment state transitions before producing the predicted observation.
Key behavior:
- Thinking mode is active by default — the model generates internal reasoning tokens enclosed in
<think>tags before the final output. - This enables the model to reason through complex environment simulation scenarios (e.g., predicting terminal output, web page state, or Android UI changes) before emitting the predicted observation.
Recommended sampling parameters when using thinking mode:
temperature=0.6top_p=0.95top_k=20
Context Length
- Native context length: 262,144 tokens (256K)
- Extended context length: 262,144 tokens (same as native)
- Recommendation: The model leverages extended context for multi-turn environment simulation. If OOM errors occur, consider reducing the context window, but maintain at least 128K tokens for optimal simulation performance.
Architecture note: The Hybrid Gated DeltaNet + Gated Attention architecture with 40 layers supports this context length natively. Rotary Position Embedding (RoPE) dimension is 64.
Training Pipeline
Three-stage training pipeline:
-
Continual Pre-Training (CPT): Environment knowledge injection. Base model: Qwen3.5-35B-A3B-Base. Environment modeling is the training objective from the CPT stage onward (native world model, not post-hoc adaptation).
-
Supervised Fine-Tuning (SFT): Next-state-prediction reasoning activation. The model learns to predict the next environment state given an agent's action and interaction history.
-
Reinforcement Learning (RL): GSPO (Group Sequence Policy Optimization) sharpens simulation fidelity. LWM RL warm-up on single-turn, non-agentic trajectories transfers to multi-turn, tool-calling agentic tasks.
Key constraint: No outputs from external API services are included in the training pipeline.
Model Tree and Derivatives
Base model: Qwen/Qwen3.5-35B-A3B-Base
Derived models:
- Adapters: 1 model
- Finetunes: 9 models
- Merges: 4 models
- Quantizations: 73 models (llama.cpp, LM Studio, Jan, Ollama)
Collection: Qwen-AgentWorld — 3 items
Spaces using this model: 6
Citation
If you find our work helpful, feel free to give us a cite.
@article{zuo2026qwen,
title={Qwen-agentworld: language world models for general agents},
author={Zuo, Yuxin and Xiao, Zikai and Sheng, Li and Huang, Fei and Tu, Jianhong and Liu, Yuxuan and Tang, Tianyi and Hu, Xiaomeng and Su, Yang and Lan, Qingfeng and others},
journal={arXiv preprint arXiv:2606.24597},
year={2026}
}
Links: 📑 Technical Report | 📖 Blog | 🤗 Hugging Face | 🤖 ModelScope | 💻 GitHub | 🖥️ Demo
AgentWorldBench Evaluation Setup
AgentWorldBench evaluates language world models by scoring each predicted environment observation on 5 dimensions: Format, Factuality, Consistency, Realism, and Quality.
Setup
## Clone the evaluation repository
git clone https://github.com/QwenLM/Qwen-AgentWorld.git
cd Qwen-AgentWorld
## Download the benchmark
huggingface-cli download Qwen/AgentWorldBench --repo-type dataset --local-dir ./AgentWorldBench
## Install dependencies
pip install openai
Run Evaluation
The evaluation follows a three-step pipeline:
cd eval
## Step 1: Run world model inference
python eval.py infer \
--data-dir ../AgentWorldBench \
--model-base-url http://localhost:8000/v1 \
--model-name Qwen/Qwen-AgentWorld-35B-A3B \
--output-dir ./results
## Step 2: Run LLM judge scoring
export OPENAI_API_KEY="your-api-key"
python eval.py judge \
--predictions ./results/predictions.jsonl \
--judge-base-url https://api.openai.com/v1 \
--judge-model gpt-5.2-2025-12-11 \
--output-dir ./results
## Step 3: Aggregate and display scores
python eval.py score --predictions ./results/judged.jsonl
Best Practices
-
Sampling Parameters: We recommend
temperature=0.6,top_p=0.95,top_k=20for world model inference. The model uses thinking mode by default (<think>...</think>) to reason about environment state transitions before producing the predicted observation. -
Adequate Output Length: We recommend an output length of 32,768 tokens for most queries. For long, multi-step trajectories, you may increase the max output length to accommodate detailed environment observations.
-
Domain-Specific System Prompts: For optimal simulation fidelity, use the domain-specific system prompts provided in the
prompts/directory of the GitHub repository.
Using via the Chat Completions API
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY",
)
## Terminal domain example
messages = [
{
"role": "system",
"content": "You are a language world model simulating a Linux terminal environment. "
"Given the user's command, predict the terminal output."
},
{
"role": "user",
"content": "Action: execute_bash\nCommand: ls -la /home/user/project/"
}
]
response = client.chat.completions.create(
model="Qwen/Qwen-AgentWorld-35B-A3B",
messages=messages,
max_tokens=32768,
temperature=0.6,
)
print(response.choices[0].message.content)
Domain-specific world model system prompt templates are provided in
prompts/of the GitHub repository for all 7 domains. Each domain folder contains asystem_prompt.txt(world model system prompt) and ajudge_system_prompt.txt(evaluation prompt).
Inference with Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen-AgentWorld-35B-A3B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto",
)
messages = [
{
"role": "system",
"content": "You are a language world model simulating a Linux terminal environment. "
"Given the user's command, predict the terminal output."
},
{
"role": "user",
"content": "Action: execute_bash\nCommand: ls -la /home/user/project/"
}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=2048, temperature=0.6)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
print(response)
Serving with vLLM
vLLM is a high-throughput and memory-efficient inference engine for LLMs.
vllm serve Qwen/Qwen-AgentWorld-35B-A3B \
--port 8000 \
--tensor-parallel-size 4 \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--language-model-only \
--trust-remote-code
The
--language-model-onlyflag is required because the model architecture includes visual component definitions but the checkpoint only contains language model weights. Without this flag, vLLM will attempt to initialize visual modules and fail.
An OpenAI-compatible API will be available at http://localhost:8000/v1.
Serving with SGLang
SGLang is a fast serving framework for large language models.
python -m sglang.launch_server \
--model-path Qwen/Qwen-AgentWorld-35B-A3B \
--port 8000 \
--tp-size 4 \
--context-length 262144 \
--reasoning-parser qwen3
An OpenAI-compatible API will be available at http://localhost:8000/v1.
AgentWorldBench — Open-Ended Evaluation
Five-dimensional rubric mean per domain, normalized to 0-100 scale.
| Model | MCP | Search | Term. | SWE | Android | Web | OS | Overall |
|---|---|---|---|---|---|---|---|---|
| GPT-5.4 | 70.10 | 37.26 | 53.69 | 66.29 | 60.00 | 51.80 | 68.58 | 58.25 |
| Claude Opus 4.8 | 54.93 | 35.14 | 59.18 | 64.10 | 61.50 | 54.66 | 66.62 | 56.59 |
| Claude Opus 4.6 | 69.90 | 29.30 | 57.51 | 64.55 | 61.74 | 51.42 | 70.20 | 57.80 |
| Gemini 3.1 Pro | 59.07 | 30.21 | 52.47 | 59.07 | 61.40 | 52.83 | 66.92 | 54.57 |
| Claude Sonnet 4.6 | 70.00 | 28.79 | 56.98 | 64.52 | 58.03 | 50.78 | 63.17 | 56.04 |
| DeepSeek-V4-Pro | 63.27 | 27.61 | 51.26 | 59.44 | 55.17 | 50.32 | 63.70 | 52.97 |
| GLM-5.1 | 67.60 | 22.46 | 47.32 | 52.07 | 59.10 | 51.50 | 59.13 | 51.31 |
| Kimi K2.6 | 65.23 | 27.48 | 52.54 | 58.77 | 58.93 | 50.20 | 60.80 | 53.42 |
| MiniMax-M2.7 | 55.82 | 27.30 | 41.62 | 37.44 | 52.40 | 50.52 | 57.73 | 46.12 |
| Qwen3.5-35B-A3B | 57.87 | 25.98 | 46.13 | 47.58 | 53.18 | 47.10 | 56.27 | 47.73 |
| Qwen3.5-397B-A17B | 68.31 | 30.81 | 55.30 | 64.44 | 54.90 | 48.55 | 60.85 | 54.74 |
| Qwen3.6-Plus | 55.28 | 21.94 | 50.58 | 59.08 | 57.65 | 50.78 | 60.33 | 50.81 |
| Qwen-AgentWorld-35B-A3B | 64.79 | 36.69 | 53.96 | 65.63 | 58.17 | 49.55 | 65.92 | 56.39 |
| Qwen-AgentWorld-397B-A17B | 68.24 | 37.82 | 57.73 | 68.49 | 60.20 | 50.98 | 67.89 | 58.71 |
Source: HuggingFace Model Card
Model Overview — Architecture Details
- Type: Causal Language Model (Language World Model)
- Base Model: Qwen3.5-35B-A3B-Base
- Training Stage: Continual Pre-Training (CPT) → Supervised Fine-Tuning (SFT) → Reinforcement Learning (RL, GSPO)
- Number of Parameters: 35B in total and 3B activated
- Hidden Dimension: 2048
- Token Embedding: 248320 (Padded)
- Number of Layers: 40
- Hidden Layout: 10 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE))
Gated DeltaNet:
- Number of Linear Attention Heads: 32 for V and 16 for QK
- Head Dimension: 128
Gated Attention:
- Number of Attention Heads: 16 for Q and 2 for KV
- Head Dimension: 256
- Rotary Position Embedding Dimension: 64
Mixture of Experts:
-
Number of Experts: 256
-
Number of Activated Experts: 8 Routed + 1 Shared
-
Expert Intermediate Dimension: 512
-
Context Length: 262,144 tokens
-
Disclaimer: No outputs from external API services are included in the training pipeline.
Key Highlights
- Seven Unified Domains. A single model covers MCP (tool calling), Search, Terminal, SWE (software engineering), Android, Web, and OS — spanning both text and GUI interaction environments.
- Native World Model. Environment modeling from CPT onward, not post-hoc adaptation on a general-purpose LLM.
- Generalizable, Scalable & Controllable Simulator. Zero-shot generalization to OOD environments (e.g., OpenClaw); controllable perturbations and fictional-world construction surpass real-environment training.
- Agent Foundation Model. LWM RL warm-up on single-turn, non-agentic trajectories transfers to multi-turn, tool-calling agentic tasks across 7 benchmarks, including 3 entirely out-of-domain.
Model Description
Qwen-AgentWorld is the first language world model to cover seven agent interaction domains within a single model. It simulates agentic environments via long chain-of-thought reasoning, predicting the next environment state given an agent's action and interaction history.
Trained through a three-stage pipeline — CPT injects environment knowledge, SFT activates next-state-prediction reasoning, RL sharpens simulation fidelity — Qwen-AgentWorld is a native world model: environment modeling is the training objective from the CPT stage onward, not a post-hoc add-on.
This repository contains the model weights and configuration files for Qwen-AgentWorld-35B-A3B, a native language world model trained for agentic environment simulation. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc.
Architecture
- Attention
- Hybrid Attention (16:2)
- MoE
- 256 experts · top-8 per token
- Layers
- 40
- Hidden size
- 2048
- Context
- 262K tokens
- Parameters
- 35000M
- Active params
- 3000M
Source: Hugging Face config.json · Qwen3_5MoeForConditionalGeneration · exact layer pattern · model repo
248K
Training Pipeline
-
1
cpt
Continual Pre-Training - Environment Knowledge Injection
CPT stage injects environment knowledge into the base model (Qwen3.5-35B-A3B-Base). Environment modeling is the training objective from CPT onward, making Qwen-AgentWorld a native world model rather than a post-hoc adaptation.
-
2
sft
Supervised Fine-Tuning - Next-State-Prediction Reasoning
SFT stage activates next-state-prediction reasoning. The model learns to predict the next environment state given an agent's action and interaction history via long chain-of-thought reasoning.
-
3
rl
Reinforcement Learning - Simulation Fidelity (GSPO)
RL stage using GSPO sharpens simulation fidelity. LWM RL warm-up on single-turn non-agentic trajectories transfers to multi-turn tool-calling agentic tasks across 7 benchmarks.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| Environment knowledge corpus (CPT injection) | pretraining | — | — | |
| Next-state-prediction SFT + GSPO RL agent trajectories | rl | — | — |
Linked Resources
Qwen-AgentWorld: Language World Models for General Agents (arXiv:2606.24597)
http://arxiv.org/abs/2606.24597
Qwen-AgentWorld Blog Post
https://qwen.ai/blog?id=qwen-agentworld
QwenLM/Qwen-AgentWorld GitHub Repository
https://github.com/QwenLM/Qwen-AgentWorld
Qwen-AgentWorld HuggingFace Collection
https://huggingface.co/collections/Qwen/qwen-agentworld
Qwen-AgentWorld on ModelScope
https://modelscope.cn/collections/Qwen/Qwen-AgentWorld
Qwen-AgentWorld Interactive Demo
https://qwen.ai/blog?id=qwen-agentworld#interactive-demo-interactive-demo
BibTeX Citation for Qwen-AgentWorld
https://arxiv.org/abs/2606.24597
Trend Analysis
24h Change
+0.2%
7d Change
+1.9%
Current
101,994
downloads
-1.9%
likes
+0.1%
downloads_all_time
+0.7%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| huggingface | downloads_all_time | 212,965 | daily | 01.09.2026 |
| huggingface | followers | 101,994 | daily | 01.09.2026 |
| huggingface | likes | 710 | daily | 01.09.2026 |
| huggingface | downloads | 71,245 | daily | 01.09.2026 |
| huggingface | downloads_all_time | 211,399 | daily | 31.08.2026 |
| huggingface | followers | 101,772 | daily | 31.08.2026 |
| huggingface | likes | 709 | daily | 31.08.2026 |
| huggingface | downloads | 72,612 | daily | 31.08.2026 |
| huggingface | downloads_all_time | 210,222 | daily | 30.08.2026 |
| huggingface | followers | 101,528 | daily | 30.08.2026 |
| huggingface | likes | 710 | daily | 30.08.2026 |
| huggingface | downloads | 74,710 | daily | 30.08.2026 |
| huggingface | followers | 101,306 | daily | 29.08.2026 |
| huggingface | likes | 709 | daily | 29.08.2026 |
| huggingface | downloads | 74,836 | daily | 29.08.2026 |
| huggingface | followers | 101,117 | daily | 28.08.2026 |
| huggingface | likes | 709 | daily | 28.08.2026 |
| huggingface | downloads | 74,645 | daily | 28.08.2026 |
| huggingface | followers | 100,862 | daily | 27.08.2026 |
| huggingface | likes | 707 | daily | 27.08.2026 |
| huggingface | downloads | 77,162 | daily | 27.08.2026 |
| huggingface | followers | 100,537 | daily | 26.08.2026 |
| huggingface | likes | 707 | daily | 26.08.2026 |
| huggingface | downloads | 78,938 | daily | 26.08.2026 |
| huggingface | followers | 100,133 | daily | 25.08.2026 |
| huggingface | likes | 704 | daily | 25.08.2026 |
| huggingface | downloads | 79,414 | daily | 25.08.2026 |
| huggingface | followers | 99,896 | daily | 24.08.2026 |
| huggingface | likes | 702 | daily | 24.08.2026 |
| huggingface | downloads | 79,340 | daily | 24.08.2026 |
| huggingface | followers | 99,669 | daily | 23.08.2026 |
| huggingface | likes | 700 | daily | 23.08.2026 |
| huggingface | downloads | 79,347 | daily | 23.08.2026 |
| huggingface | followers | 99,461 | daily | 22.08.2026 |
| huggingface | likes | 698 | daily | 22.08.2026 |
| huggingface | downloads | 79,709 | daily | 22.08.2026 |
| huggingface | followers | 99,274 | daily | 21.08.2026 |
| huggingface | likes | 697 | daily | 21.08.2026 |
| huggingface | downloads | 81,037 | daily | 21.08.2026 |
| huggingface | followers | 99,037 | daily | 20.08.2026 |
| huggingface | likes | 695 | daily | 20.08.2026 |
| huggingface | downloads | 78,996 | daily | 20.08.2026 |
| huggingface | followers | 98,772 | daily | 19.08.2026 |
| huggingface | likes | 693 | daily | 19.08.2026 |
| huggingface | downloads | 79,219 | daily | 19.08.2026 |
| huggingface | followers | 98,465 | daily | 18.08.2026 |
| huggingface | likes | 691 | daily | 18.08.2026 |
| huggingface | downloads | 80,074 | daily | 18.08.2026 |
| huggingface | likes | 356 likes | snapshot | 27.06.2026 |
| huggingface | downloads | 18,872 downloads | monthly | 27.06.2026 |