Qwen-AgentWorld-35B-A3B

Qwen Team (Alibaba Cloud)

Parameters

35.0B total / 3.0B active

MoE: total / active

Architecture

Hybrid Gated DeltaNet + Gated Attention MoE - Language World Model (qwen3_5_moe)

Released

22.06.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16 Qwen-AgentWorld en zh

Input Modalities

text image

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

262,144 tokens

Openness Index Score 100.0/100

About

Qwen-AgentWorld-35B-A3B (Qwen/Qwen-AgentWorld-35B-A3B) is Alibaba's open Language World Model (LWM) - an agent foundation model built as a native world model rather than a post-hoc-adapted general LLM. A single model covers seven unified domains - MCP (tool calling), Search, Terminal, SWE, Android, Web and OS - spanning both text and GUI interaction environments. It generalizes zero-shot to out-of-distribution environments (e.g. OpenClaw), supports controllable perturbations, and can construct fictional training worlds that surpass real-environment training.

Built on Qwen3.5-35B-A3B-Base (35B total, 3B activated; 40 layers in a 10x (3x (Gated DeltaNet -> MoE) + 1x (Gated Attention -> MoE)) hybrid layout with 256 experts, 8 routed + 1 shared per token, hidden 2048), it is trained with a three-stage pipeline: Continual Pre-Training for environment knowledge injection, SFT with next-state-prediction reasoning, and RL with GSPO for simulation fidelity. LWM RL warm-up on single-turn, non-agentic trajectories transfers to multi-turn, tool-calling agentic tasks across 7 benchmarks including 3 entirely out-of-domain. A key disclaimer: no outputs from external API services are included in the training pipeline. Context length is 262,144 tokens; released June 22, 2026 under Apache 2.0. Paper: "Qwen-AgentWorld: Language World Models for General Agents" (arXiv 2606.24597).

Training Data Three-stage pipeline: CPT (environment knowledge injection) → SFT (next-state-prediction reasoning) → RL (GSPO, simulation fidelity). No external API service outputs included in training. Base model: Qwen3.5-35B-A3B-Base.

Benchmark Scores

Benchmark Score Date
AgentWorldBench MCP
general_agent
65.00%
24.06.2026
AgentWorldBench Search
general_agent
92.88%
24.06.2026
AgentWorldBench Terminal
general_agent
70.27%
24.06.2026
AgentWorldBench SWE
general_agent
90.79%
24.06.2026
AgentWorldBench Web
general_agent
32.41%
24.06.2026
AgentWorldBench OS
general_agent
69.27%
24.06.2026
AgentWorldBench Overall
general_agent
81.57%
24.06.2026
AgentWorldBench Android
general_agent
61.78%
24.06.2026

Thinking Mode

The model uses thinking mode by default (<think>...</think>) to reason about environment state transitions before producing the predicted observation.

Key behavior:

  • Thinking mode is active by default — the model generates internal reasoning tokens enclosed in <think> tags before the final output.
  • This enables the model to reason through complex environment simulation scenarios (e.g., predicting terminal output, web page state, or Android UI changes) before emitting the predicted observation.

Recommended sampling parameters when using thinking mode:

  • temperature=0.6
  • top_p=0.95
  • top_k=20

Context Length

  • Native context length: 262,144 tokens (256K)
  • Extended context length: 262,144 tokens (same as native)
  • Recommendation: The model leverages extended context for multi-turn environment simulation. If OOM errors occur, consider reducing the context window, but maintain at least 128K tokens for optimal simulation performance.

Architecture note: The Hybrid Gated DeltaNet + Gated Attention architecture with 40 layers supports this context length natively. Rotary Position Embedding (RoPE) dimension is 64.

Training Pipeline

Three-stage training pipeline:

  1. Continual Pre-Training (CPT): Environment knowledge injection. Base model: Qwen3.5-35B-A3B-Base. Environment modeling is the training objective from the CPT stage onward (native world model, not post-hoc adaptation).

  2. Supervised Fine-Tuning (SFT): Next-state-prediction reasoning activation. The model learns to predict the next environment state given an agent's action and interaction history.

  3. Reinforcement Learning (RL): GSPO (Group Sequence Policy Optimization) sharpens simulation fidelity. LWM RL warm-up on single-turn, non-agentic trajectories transfers to multi-turn, tool-calling agentic tasks.

Key constraint: No outputs from external API services are included in the training pipeline.

Model Tree and Derivatives

Base model: Qwen/Qwen3.5-35B-A3B-Base

Derived models:

  • Adapters: 1 model
  • Finetunes: 9 models
  • Merges: 4 models
  • Quantizations: 73 models (llama.cpp, LM Studio, Jan, Ollama)

Collection: Qwen-AgentWorld — 3 items

Spaces using this model: 6

Citation

If you find our work helpful, feel free to give us a cite.

@article{zuo2026qwen,
  title={Qwen-agentworld: language world models for general agents},
  author={Zuo, Yuxin and Xiao, Zikai and Sheng, Li and Huang, Fei and Tu, Jianhong and Liu, Yuxuan and Tang, Tianyi and Hu, Xiaomeng and Su, Yang and Lan, Qingfeng and others},
  journal={arXiv preprint arXiv:2606.24597},
  year={2026}
}

Links: 📑 Technical Report | 📖 Blog | 🤗 Hugging Face | 🤖 ModelScope | 💻 GitHub | 🖥️ Demo

AgentWorldBench Evaluation Setup

AgentWorldBench evaluates language world models by scoring each predicted environment observation on 5 dimensions: Format, Factuality, Consistency, Realism, and Quality.

Setup

## Clone the evaluation repository
git clone https://github.com/QwenLM/Qwen-AgentWorld.git
cd Qwen-AgentWorld

## Download the benchmark
huggingface-cli download Qwen/AgentWorldBench --repo-type dataset --local-dir ./AgentWorldBench

## Install dependencies
pip install openai

Run Evaluation

The evaluation follows a three-step pipeline:

cd eval

## Step 1: Run world model inference
python eval.py infer \
    --data-dir ../AgentWorldBench \
    --model-base-url http://localhost:8000/v1 \
    --model-name Qwen/Qwen-AgentWorld-35B-A3B \
    --output-dir ./results

## Step 2: Run LLM judge scoring
export OPENAI_API_KEY="your-api-key"
python eval.py judge \
    --predictions ./results/predictions.jsonl \
    --judge-base-url https://api.openai.com/v1 \
    --judge-model gpt-5.2-2025-12-11 \
    --output-dir ./results

## Step 3: Aggregate and display scores
python eval.py score --predictions ./results/judged.jsonl

Best Practices

  1. Sampling Parameters: We recommend temperature=0.6, top_p=0.95, top_k=20 for world model inference. The model uses thinking mode by default (<think>...</think>) to reason about environment state transitions before producing the predicted observation.

  2. Adequate Output Length: We recommend an output length of 32,768 tokens for most queries. For long, multi-step trajectories, you may increase the max output length to accommodate detailed environment observations.

  3. Domain-Specific System Prompts: For optimal simulation fidelity, use the domain-specific system prompts provided in the prompts/ directory of the GitHub repository.

Using via the Chat Completions API

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="EMPTY",
)

## Terminal domain example
messages = [
    {
        "role": "system",
        "content": "You are a language world model simulating a Linux terminal environment. "
                   "Given the user's command, predict the terminal output."
    },
    {
        "role": "user",
        "content": "Action: execute_bash\nCommand: ls -la /home/user/project/"
    }
]

response = client.chat.completions.create(
    model="Qwen/Qwen-AgentWorld-35B-A3B",
    messages=messages,
    max_tokens=32768,
    temperature=0.6,
)
print(response.choices[0].message.content)

Domain-specific world model system prompt templates are provided in prompts/ of the GitHub repository for all 7 domains. Each domain folder contains a system_prompt.txt (world model system prompt) and a judge_system_prompt.txt (evaluation prompt).

Inference with Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen/Qwen-AgentWorld-35B-A3B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto",
)

messages = [
    {
        "role": "system",
        "content": "You are a language world model simulating a Linux terminal environment. "
                   "Given the user's command, predict the terminal output."
    },
    {
        "role": "user",
        "content": "Action: execute_bash\nCommand: ls -la /home/user/project/"
    }
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=2048, temperature=0.6)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
print(response)

Serving with vLLM

vLLM is a high-throughput and memory-efficient inference engine for LLMs.

vllm serve Qwen/Qwen-AgentWorld-35B-A3B \
    --port 8000 \
    --tensor-parallel-size 4 \
    --max-model-len 262144 \
    --reasoning-parser qwen3 \
    --language-model-only \
    --trust-remote-code

The --language-model-only flag is required because the model architecture includes visual component definitions but the checkpoint only contains language model weights. Without this flag, vLLM will attempt to initialize visual modules and fail.

An OpenAI-compatible API will be available at http://localhost:8000/v1.

Serving with SGLang

SGLang is a fast serving framework for large language models.

python -m sglang.launch_server \
    --model-path Qwen/Qwen-AgentWorld-35B-A3B \
    --port 8000 \
    --tp-size 4 \
    --context-length 262144 \
    --reasoning-parser qwen3

An OpenAI-compatible API will be available at http://localhost:8000/v1.

AgentWorldBench — Open-Ended Evaluation

Five-dimensional rubric mean per domain, normalized to 0-100 scale.

Model MCP Search Term. SWE Android Web OS Overall
GPT-5.4 70.10 37.26 53.69 66.29 60.00 51.80 68.58 58.25
Claude Opus 4.8 54.93 35.14 59.18 64.10 61.50 54.66 66.62 56.59
Claude Opus 4.6 69.90 29.30 57.51 64.55 61.74 51.42 70.20 57.80
Gemini 3.1 Pro 59.07 30.21 52.47 59.07 61.40 52.83 66.92 54.57
Claude Sonnet 4.6 70.00 28.79 56.98 64.52 58.03 50.78 63.17 56.04
DeepSeek-V4-Pro 63.27 27.61 51.26 59.44 55.17 50.32 63.70 52.97
GLM-5.1 67.60 22.46 47.32 52.07 59.10 51.50 59.13 51.31
Kimi K2.6 65.23 27.48 52.54 58.77 58.93 50.20 60.80 53.42
MiniMax-M2.7 55.82 27.30 41.62 37.44 52.40 50.52 57.73 46.12
Qwen3.5-35B-A3B 57.87 25.98 46.13 47.58 53.18 47.10 56.27 47.73
Qwen3.5-397B-A17B 68.31 30.81 55.30 64.44 54.90 48.55 60.85 54.74
Qwen3.6-Plus 55.28 21.94 50.58 59.08 57.65 50.78 60.33 50.81
Qwen-AgentWorld-35B-A3B 64.79 36.69 53.96 65.63 58.17 49.55 65.92 56.39
Qwen-AgentWorld-397B-A17B 68.24 37.82 57.73 68.49 60.20 50.98 67.89 58.71

Source: HuggingFace Model Card

Model Overview — Architecture Details

  • Type: Causal Language Model (Language World Model)
  • Base Model: Qwen3.5-35B-A3B-Base
  • Training Stage: Continual Pre-Training (CPT) → Supervised Fine-Tuning (SFT) → Reinforcement Learning (RL, GSPO)
  • Number of Parameters: 35B in total and 3B activated
  • Hidden Dimension: 2048
  • Token Embedding: 248320 (Padded)
  • Number of Layers: 40
  • Hidden Layout: 10 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE))

Gated DeltaNet:

  • Number of Linear Attention Heads: 32 for V and 16 for QK
  • Head Dimension: 128

Gated Attention:

  • Number of Attention Heads: 16 for Q and 2 for KV
  • Head Dimension: 256
  • Rotary Position Embedding Dimension: 64

Mixture of Experts:

  • Number of Experts: 256

  • Number of Activated Experts: 8 Routed + 1 Shared

  • Expert Intermediate Dimension: 512

  • Context Length: 262,144 tokens

  • Disclaimer: No outputs from external API services are included in the training pipeline.

Key Highlights

  • Seven Unified Domains. A single model covers MCP (tool calling), Search, Terminal, SWE (software engineering), Android, Web, and OS — spanning both text and GUI interaction environments.
  • Native World Model. Environment modeling from CPT onward, not post-hoc adaptation on a general-purpose LLM.
  • Generalizable, Scalable & Controllable Simulator. Zero-shot generalization to OOD environments (e.g., OpenClaw); controllable perturbations and fictional-world construction surpass real-environment training.
  • Agent Foundation Model. LWM RL warm-up on single-turn, non-agentic trajectories transfers to multi-turn, tool-calling agentic tasks across 7 benchmarks, including 3 entirely out-of-domain.

Model Description

Qwen-AgentWorld is the first language world model to cover seven agent interaction domains within a single model. It simulates agentic environments via long chain-of-thought reasoning, predicting the next environment state given an agent's action and interaction history.

Trained through a three-stage pipeline — CPT injects environment knowledge, SFT activates next-state-prediction reasoning, RL sharpens simulation fidelity — Qwen-AgentWorld is a native world model: environment modeling is the training objective from the CPT stage onward, not a post-hoc add-on.

This repository contains the model weights and configuration files for Qwen-AgentWorld-35B-A3B, a native language world model trained for agentic environment simulation. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc.

Architecture

Decoder Block ×40 input Embedding vocab 248K · d 2048 Linear / Recurrent Hybrid 16:2 · dₕ 256 ×30 Full Attention Hybrid 16:2 · dₕ 256 ×10 MoE FFN 256 experts · top-8 · dᴻ 512 Final Norm LM Head vocab 248K output
Attention
Hybrid Attention (16:2)
MoE
256 experts · top-8 per token
Layers
40
Hidden size
2048
Context
262K tokens
Parameters
35000M
Active params
3000M

Source: Hugging Face config.json · Qwen3_5MoeForConditionalGeneration · exact layer pattern · model repo

Type: Hybrid Gated DeltaNet + Gated Attention MoE
Attention: 10 x (3 x Gated DeltaNet then 1 x Gated Attention) with MoE throughout
Decoder: Causal Language Model - Language World Model
MoE: yes (256 experts)
Routing: Top-K routing with 8 routed experts plus 1 shared expert
Layers 40
Context length 262K
Experts 256
Experts per token 9
Hidden size 2048
Expert FFN dim 512
Vision Yes
Gated Attention Head Dim 256
Gated Attention Heads Kv 2
Gated Attention Heads Q 16
Gated Deltanet Head Dim 128
Gated Deltanet Heads Qk 16
Gated Deltanet Heads V 32
RoPE dim 64
Token Embedding

248K

Training Pipeline

  1. 1
    cpt

    Continual Pre-Training - Environment Knowledge Injection

    CPT stage injects environment knowledge into the base model (Qwen3.5-35B-A3B-Base). Environment modeling is the training objective from CPT onward, making Qwen-AgentWorld a native world model rather than a post-hoc adaptation.

  2. 2
    sft

    Supervised Fine-Tuning - Next-State-Prediction Reasoning

    SFT stage activates next-state-prediction reasoning. The model learns to predict the next environment state given an agent's action and interaction history via long chain-of-thought reasoning.

  3. 3
    rl

    Reinforcement Learning - Simulation Fidelity (GSPO)

    RL stage using GSPO sharpens simulation fidelity. LWM RL warm-up on single-turn non-agentic trajectories transfers to multi-turn tool-calling agentic tasks across 7 benchmarks.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Environment knowledge corpus (CPT injection) pretraining — —
Next-state-prediction SFT + GSPO RL agent trajectories rl — —

Trend Analysis

24h Change

+0.2%

7d Change

+1.9%

Current

101,994

huggingface

downloads

-1.9%

huggingface

likes

+0.1%

huggingface

downloads_all_time

+0.7%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
huggingface downloads_all_time 212,965 daily 01.09.2026
huggingface followers 101,994 daily 01.09.2026
huggingface likes 710 daily 01.09.2026
huggingface downloads 71,245 daily 01.09.2026
huggingface downloads_all_time 211,399 daily 31.08.2026
huggingface followers 101,772 daily 31.08.2026
huggingface likes 709 daily 31.08.2026
huggingface downloads 72,612 daily 31.08.2026
huggingface downloads_all_time 210,222 daily 30.08.2026
huggingface followers 101,528 daily 30.08.2026
huggingface likes 710 daily 30.08.2026
huggingface downloads 74,710 daily 30.08.2026
huggingface followers 101,306 daily 29.08.2026
huggingface likes 709 daily 29.08.2026
huggingface downloads 74,836 daily 29.08.2026
huggingface followers 101,117 daily 28.08.2026
huggingface likes 709 daily 28.08.2026
huggingface downloads 74,645 daily 28.08.2026
huggingface followers 100,862 daily 27.08.2026
huggingface likes 707 daily 27.08.2026
huggingface downloads 77,162 daily 27.08.2026
huggingface followers 100,537 daily 26.08.2026
huggingface likes 707 daily 26.08.2026
huggingface downloads 78,938 daily 26.08.2026
huggingface followers 100,133 daily 25.08.2026
huggingface likes 704 daily 25.08.2026
huggingface downloads 79,414 daily 25.08.2026
huggingface followers 99,896 daily 24.08.2026
huggingface likes 702 daily 24.08.2026
huggingface downloads 79,340 daily 24.08.2026
huggingface followers 99,669 daily 23.08.2026
huggingface likes 700 daily 23.08.2026
huggingface downloads 79,347 daily 23.08.2026
huggingface followers 99,461 daily 22.08.2026
huggingface likes 698 daily 22.08.2026
huggingface downloads 79,709 daily 22.08.2026
huggingface followers 99,274 daily 21.08.2026
huggingface likes 697 daily 21.08.2026
huggingface downloads 81,037 daily 21.08.2026
huggingface followers 99,037 daily 20.08.2026
huggingface likes 695 daily 20.08.2026
huggingface downloads 78,996 daily 20.08.2026
huggingface followers 98,772 daily 19.08.2026
huggingface likes 693 daily 19.08.2026
huggingface downloads 79,219 daily 19.08.2026
huggingface followers 98,465 daily 18.08.2026
huggingface likes 691 daily 18.08.2026
huggingface downloads 80,074 daily 18.08.2026
huggingface likes 356 likes snapshot 27.06.2026
huggingface downloads 18,872 downloads monthly 27.06.2026

View full metric history →

Related Models