Spark-X2.5-4B

SparkLLM (XHToken)

Parameters

4.0B

Architecture

Hybrid attention transformer: 1 full-attention layer combined with 3 sliding-window attention (SWA) layers, natively supporting up to 1M-token context

Released

01.09.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16 Spark-X2.5 en zh multilingual (200+ languages)

Input Modalities

text

Output Modalities

text

Context (native)

1,048,576 tokens

Context (extended)

1,048,576 tokens

Openness Index Score 100.0/100

About

Spark-X2.5-4B is a compact, general-purpose text language model from SparkLLM (XHToken), open-sourced under Apache 2.0 with a native context window of up to 1M tokens.

Technical Highlights

  • Efficient Architecture and Native 1M-token Context: A hybrid attention architecture combining one full-attention layer with three sliding-window attention layers. This design substantially reduces the computational overhead typically associated with long-context models while natively supporting a context window of up to 1M tokens.
  • Strong Coding and Agent Capabilities: Deeply integrated with popular agent harnesses including Codex, Claude Code, OpenClaw, and Hermes. State-of-the-art performance among models of comparable size across everyday coding, agentic workflows, reasoning, and instruction-following tasks.
  • Broad Hardware and Software Compatibility: Supports NVIDIA, Huawei, Hygon, HOUMO.AI, and other hardware platforms. Compatible with vLLM, SGLang, llama.cpp, and MLX; deployable via Ollama and LM Studio. Customizable with LLaMA-Factory. Delivers superior TTFT, TOPT, and overall inference efficiency compared with similarly sized models.
  • Advanced Training Algorithms: Trained on Huawei Ascend clusters. Large-scale reinforcement learning and post-training techniques such as MOPD significantly enhance reasoning, coding, agentic, and instruction-following capabilities.

Trained on ~20T tokens with large-scale reinforcement learning and MOPD consolidation, it supports 200+ languages and posts leading results among similarly sized open models on agentic, coding, math, and instruction-following benchmarks. Aimed at agentic workflows, coding, reasoning, and long-context tasks, it is built for practical, efficient on-device deployment.

Training Data Pretrained on ~20 trillion tokens (web pages, books, academic publications, code, encyclopedic materials) with dedicated long-context training up to 1M sequence length; post-training: supervised fine-tuning, large-scale reinforcement learning across capability domains, and MOPD consolidation of domain-specialized teacher policies. Trained on Huawei Ascend clusters. Supports 200+ languages.

Benchmark Scores

Benchmark Score Date
TAU3-Bench
general_agent
43.13%
—
VITA-Bench
general_agent
53.89%
—
SWE-bench Pro
coding_agent
55.50%
—
SWE-bench Verified
coding_agent
47.25%
—
SWE-bench Multilingual
coding_agent
57.16%
—
AIME 26
stem_reasoning
89.16%
—
HMMT Feb 26
stem_reasoning
78.63%
—
IMOAnswerBench
stem_reasoning
73.30%
—
Humanity's Last Exam
stem_reasoning
19.32%
—
Workspace Bench
100.00%
—
SciCode
84.91%
—
Gaokao 2026
96.09%
—
BFCL-V4
general_agent
83.10%
—
TAU2-Bench
general_agent
72.84%
—
BrowseComp
general_agent
42.91%
—
IFEval
instruction_following
96.68%
—
IFBench
instruction_following
87.02%
—
AA-LCR
long_context
70.38%
—
MCPMark
general_agent
12.91%
—
GPQA
reasoning
84.24%
—
MCP-Atlas
general_agent
57.46%
—

Model artifacts and ecosystem

The repository contains model weights and configuration files for the post-trained model in Hugging Face Transformers format, distributed as Safetensors: 4B parameters with BF16 tensor type. The model uses custom code (custom_code tag), so pass --trust-remote-code when serving with vLLM. The base model is XHToken/Spark-X2.5-4B-Base; this checkpoint is a fine-tune of it. The ecosystem includes 5 public fine-tunes and 18 quantizations (usable with llama.cpp, LM Studio, Jan, and Ollama), and the models are also part of the Spark-X2.5 collection on Hugging Face.

Citation

If you find the work helpful, feel free to cite it:

@misc{sparkx2.5,
    title  = {Spark-X2.5 4B&1.7B: Pushing the Limits of Agentic Capabilities in On-Device Models},
    author = {SparkLLM Team},
    year   = {2026}
}

License

The Spark-X2.5 model series is licensed under the Apache 2.0 License (see the LICENSE file in the model repository).

Recommended usage settings

  • Sampling parameters: the recommended settings for Spark-X2.5 are temperature=1.0, top_p=0.95, and top_k=-1 (example requests also use repetition_penalty=1, presence_penalty=0, frequency_penalty=0).
  • Thinking mode: thinking is enabled by default; to disable it for a single request, set "chat_template_kwargs": {"enable_thinking": false} in the request body.
  • Long outputs: example requests use max_tokens up to 131,072 for long-form generation.
  • Context length vs device memory: serving with the full 1,048,576-token context requires sufficient device memory; reduce --context-length when necessary.
  • Fine-tuning: use LLaMA-Factory to customize the model.
  • Agent harnesses: the models are deeply integrated with popular agent harnesses including Codex, Claude Code, OpenClaw, and Hermes.

Deployment: SGLang, vLLM, MLX, Ollama, LM Studio

Spark-X2.5 runs on a wide range of hardware — NVIDIA GPUs, Huawei Ascend NPUs (A2, A3, and Ascend 950DT), Apple silicon, and Linux CPU — and is compatible with leading inference frameworks such as vLLM, SGLang, llama.cpp, and MLX, plus quick deployment through Ollama and LM Studio.

SGLang (NVIDIA GPU). Use the pre-built image that tracks the Spark-X2.5 runtime (lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1; Ascend NPU daily builds are published under quay.io/ascend/sglang). Serve an OpenAI-compatible API with:

docker run --rm -it --gpus '"device=0"' --ipc=host -p 30000:30000 \
  -v "$MODEL_PATH:/root/Spark-X2.5-4B:ro" \
  lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1 \
  python -m sglang.launch_server \
    --model-path /root/Spark-X2.5-4B \
    --served-model-name spark2.5 \
    --tool-call-parser spark25 \
    --reasoning-parser qwen3 \
    --tp-size 1 --mem-fraction-static 0.8 \
    --context-length 1048576 \
    --chat-template /root/Spark-X2.5-4B/chat_template.jinja \
    --host 0.0.0.0 --port 30000

Set MODEL_PATH to the absolute path of the local Spark-X2.5-4B checkpoint before starting the container.

vLLM (NVIDIA GPU). Use the official vllm/vllm-openai:latest image with --trust-remote-code and the model's chat_template.jinja (prefix caching optional):

docker run --rm --gpus all --ipc=host -p 30000:30000 \
  -v "$MODEL_PATH:/models/Spark-X2.5-4B:ro" \
  vllm/vllm-openai:latest \
  --model /models/Spark-X2.5-4B --port 30000 \
  --trust-remote-code --served-model-name spark25 \
  --tensor-parallel-size 1 --gpu-memory-utilization 0.7 \
  --enable-prefix-caching \
  --chat-template /models/Spark-X2.5-4B/chat_template.jinja

For Ascend NPUs, use the official quay.io/ascend/vllm-ascend images (nightly-main for A2, nightly-main-a3 for A3, nightly-main-a5 for Ascend 950DT), then install the Spark plugin inside the container:

git clone https://github.com/XHToken/Spark-plugin.git
cd Spark-plugin
uv pip install .

MLX (Spark-MLX-LLM). Runs the original Spark-X2.5 Hugging Face checkpoints locally on Apple silicon GPU, Linux CPU, and NVIDIA CUDA on Linux — no GGUF conversion required:

git clone https://github.com/XHToken/Spark-MLX-LLM.git
cd Spark-MLX-LLM
python3 -m venv .venv && source .venv/bin/activate
python -m pip install -e .   # Apple silicon (extras: '.[cpu]', '.[cuda12]', '.[cuda13]')
spark-mlx-generate --device gpu --dtype bfloat16 \
  --model XHToken/Spark-X2.5-4B \
  --prompt "What is the capital of Anhui Province?" \
  --max-tokens 512 --temp 0

Ollama and LM Studio. Build the provided llama.cpp fork (https://github.com/XHToken/llama.cpp), then either create an Ollama model from a GGUF file (./ollama create Spark-X2.5-4B -f ./Modelfile.spark and ./ollama run Spark-X2.5-4B) or copy the llama.cpp-spark build output into an LM Studio runtime directory, place the GGUF model under the LM Studio models directory, and load it (or use the lms CLI: lms load <model>, lms chat <model>).

Fine-tuning. LLaMA-Factory is recommended for fine-tuning the model.

Thinking mode control

Thinking is enabled by default by both the chat template and the Qwen3 reasoning parser (the SGLang server is started with --reasoning-parser qwen3). To disable thinking for a specific request, set "chat_template_kwargs": {"enable_thinking": false} in the request body. All reported benchmark evaluations are conducted in thinking mode.

Benchmarks vs similarly sized models

Spark-X2.5 models are evaluated against leading on-device models of similar size across a broad range of tasks, including agent, code, math, and general & knowledge benchmarks.

Benchmark Spark‑X2.5‑4B Spark‑X2.5‑1.7B Qwen3.5‑9B Qwen3.5‑4B Qwen3.5‑2B Gemma4‑12B Gemma4‑E4B Gemma4‑E2B
Agent
BFCL‑V4 65.1 46.9 66.1* 50.3* 43.6* 37.4 36.9 30.2
τ²‑bench 75.1 65.3 79.1* 79.9* 48.8* 69.0* 42.2* 24.5*
τ³‑bench 30.4 20.1 9.3 6.7 4.1 13.3 10.1 8.8
MCP‑Atlas 54.6 23.4 47.4* 40.8* 14.8 30.5* 15.0* 12.6
MCP‑Mark 14.2 2.3 13.4 12.5 – – – –
Workspace Bench 31.2 18.9 25.5 21.3 7.7 – – –
VitaBench2.0 25.2 8.3 15.6 18.2 5.2 12.4 4.8 4.4
BrowseComp 40.9 29.7 8.3 14.3 3.1 10.0 8.3 3.7
Code
SWE‑Bench Pro 44.4 10.4 33.8* 29.4* 1.9 21.9* 4.0* –
SWE‑Bench Verified 41.6 28.3 53.1* 38.8* 6.8 44.2* 14.0* –
SWE‑Bench Multilingual 53.3 23.3 43.3 27.7 5.0 32.5* – –
SciCode 34.7 18.2 32.7* 24.0 6.0 39.8 27.5 20.5
Math
Gaokao 2026 133.4 114.8 135.5 130.3 94.0 130.6 102.4 81.8
AIME 2026 90.7 69.4 88.2 83.0 30.8 82.1* 42.5* 37.5*
HMMT Feb 2026 81.2 48.4 70.8 69.7 21.5 65.6 34.2 20.5
IMO‑AnswerBench 74.2 45.4 69.8 68.5 – 57.2 26.9 22.6
General & Knowledge
IFEval 93.0 89.5 91.5* 89.8* 78.6* 94.8 45.3 34.8
IFBench 75.0 66.3 64.5 59.2 41.3* 73.5* 44.0* 22.7
AA‑LCR 56.3 24.3 63.0* 57.0* 25.6* 55.3* 34.7 18.3
HLE 12.3 6.3 14.3 8.6 2.1 13.1 3.9 2.5
GPQA 67.4 43.8 77.2 67.2 44.6 72.8 54.5 43.8

Notes:

  • * denotes reported results from publicly released model cards / papers; - denotes scores not yet available.
  • All evaluations are conducted in thinking mode. The recommended sampling parameters for Spark-X2.5 are temperature=1.0, top_p=0.95, and top_k=-1.
  • Gaokao 2026 consists of the five 2026 Chinese GAOKAO examinations (National I, National II, Beijing, Shanghai, Tianjin), each graded out of 150 points.

Training methods: pretraining data and post-training

Pretraining. Spark-X2.5 is pretrained on approximately 20 trillion tokens from a diverse corpus spanning web pages, books, academic publications, code, and encyclopedic materials. Particular attention is paid to data quality, domain coverage, and the sampling weights assigned to different data categories. Extensive data-mixture studies determine an effective balance among mathematics, logic, code, and other high-value domains, enabling the models to acquire broad general knowledge while developing stronger capabilities in complex reasoning and code generation. Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens.

Post-training. Post-training begins with supervised fine-tuning (SFT) on a carefully curated corpus, establishing robust instruction following, structured generation, and task completion while providing a stable policy initialization for reinforcement learning. Large-scale reinforcement learning is then applied across several capability domains — language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following — yielding a set of domain-specialized teacher policies whose complementary strengths are consolidated into a single deployable model through MOPD. The models were trained on Huawei Ascend clusters.

Native 1M-token context window

Spark-X2.5 natively supports a context window of up to 1M tokens (1,048,576). Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens. When serving with SGLang, start the server with --context-length 1048576; this setting requires sufficient device memory, so reduce --context-length when necessary. Example client requests generate up to 131,072 tokens (max_tokens: 131072).

Hybrid attention architecture

For agent tasks, balancing performance, inference speed, and cache usage has long been a key bottleneck limiting model performance. Spark-X2.5 systematically integrates and optimizes mature attention technologies, combining sliding-window attention (SWA) with a hybrid full-attention architecture: one full-attention layer is paired with three sliding-window attention layers. This design leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure, achieving an effective balance among performance, inference efficiency, and KV-cache size — thereby improving practicality and effectiveness in real-world deployment scenarios.

Introduction and technical highlights

Spark-X2.5-4B and Spark-X2.5-1.7B are compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. They deliver strong performance across everyday tasks — conversation, writing, translation, reasoning, coding, tool use, and agentic workflows — achieving leading results among open-source models of comparable size. They combine an efficiency-oriented architecture with native context windows of up to 1M tokens and support for more than 200 languages.

Technical highlights

  • Efficient architecture and native 1M-token context: a hybrid attention architecture combines one full-attention layer with three sliding-window attention layers, substantially reducing the computational overhead of long-context models while natively supporting a context window of up to 1M tokens.
  • Strong coding and agent capabilities: deeply integrated with popular agent harnesses including Codex, Claude Code, OpenClaw, and Hermes; state-of-the-art performance among models of comparable size across everyday coding, agentic workflows, reasoning, and instruction following.
  • Broad hardware and software compatibility: supports NVIDIA, Huawei, Hygon, HOUMO.AI, and other hardware platforms; compatible with leading inference frameworks such as vLLM, SGLang, llama.cpp, and MLX; deployable quickly through Ollama and LM Studio; customizable with fine-tuning frameworks such as LLaMA-Factory. Delivers superior TTFT, TOPT, and overall inference efficiency compared with similarly sized models.
  • Advanced training algorithms: trained on Huawei Ascend clusters; large-scale reinforcement learning and post-training techniques such as MOPD significantly enhance reasoning, coding, agentic, and instruction-following capabilities.

Architecture

Decoder Block ×36 input Embedding vocab 131K · d 2560 Sliding Window Attn Hybrid 16:4 · dₕ 256 · win 512 ×27 Full Attention Hybrid 16:4 · dₕ 256 · win 512 ×9 Dense FFN gelu · d 10K Final Norm LM Head tied with embedding output
Attention
Hybrid Attention (16:4)
Layers
36
Hidden size
2560
Context
1M tokens
Parameters
4000M

Source: Hugging Face config.json · Spark2_5ForCausalLM · exact layer pattern · model repo

Type: hybrid attention transformer (full attention + sliding-window attention)
Attention: hybrid: one full-attention layer interleaved with three sliding-window attention (SWA) layers
Context length 1M
Full Attention Layers Per Group 1
Swa Layers Per Group 3

Training Pipeline

  1. 1
    pretraining

    Pretraining (~20T tokens)

    Pretrained on approximately 20 trillion tokens from web pages, books, academic publications, code, and encyclopedic materials, with data-mixture studies balancing mathematics, logic, and code; long-context capability developed in a dedicated stage with sequence lengths up to 1M tokens.

  2. 2
    sft

    Supervised fine-tuning

    Supervised fine-tuning on a carefully curated corpus establishing instruction following, structured generation, and task completion, providing a stable policy initialization for reinforcement learning.

  3. 3
    rl

    Large-scale reinforcement learning

    Large-scale RL across capability domains including language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following, yielding domain-specialized teacher policies.

  4. 4
    other

    MOPD policy consolidation

    MOPD consolidates the complementary strengths of the domain-specialized teacher policies into a single deployable model.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
~20T token pre-training corpus (web, books, academic, code, encyclopedic) pretraining — —
SFT + large-scale RL + MOPD consolidation corpora rl — —

Related Models