Parameters
4.0B
Architecture
Hybrid attention transformer: 1 full-attention layer combined with 3 sliding-window attention (SWA) layers, natively supporting up to 1M-token context
Released
01.09.2026
License
Apache License 2.0
Input Modalities
Output Modalities
Context (native)
1,048,576 tokens
Context (extended)
1,048,576 tokens
About
Spark-X2.5-4B is a compact, general-purpose text language model from SparkLLM (XHToken), open-sourced under Apache 2.0 with a native context window of up to 1M tokens.
Technical Highlights
- Efficient Architecture and Native 1M-token Context: A hybrid attention architecture combining one full-attention layer with three sliding-window attention layers. This design substantially reduces the computational overhead typically associated with long-context models while natively supporting a context window of up to 1M tokens.
- Strong Coding and Agent Capabilities: Deeply integrated with popular agent harnesses including Codex, Claude Code, OpenClaw, and Hermes. State-of-the-art performance among models of comparable size across everyday coding, agentic workflows, reasoning, and instruction-following tasks.
- Broad Hardware and Software Compatibility: Supports NVIDIA, Huawei, Hygon, HOUMO.AI, and other hardware platforms. Compatible with vLLM, SGLang, llama.cpp, and MLX; deployable via Ollama and LM Studio. Customizable with LLaMA-Factory. Delivers superior TTFT, TOPT, and overall inference efficiency compared with similarly sized models.
- Advanced Training Algorithms: Trained on Huawei Ascend clusters. Large-scale reinforcement learning and post-training techniques such as MOPD significantly enhance reasoning, coding, agentic, and instruction-following capabilities.
Trained on ~20T tokens with large-scale reinforcement learning and MOPD consolidation, it supports 200+ languages and posts leading results among similarly sized open models on agentic, coding, math, and instruction-following benchmarks. Aimed at agentic workflows, coding, reasoning, and long-context tasks, it is built for practical, efficient on-device deployment.
Training Data Pretrained on ~20 trillion tokens (web pages, books, academic publications, code, encyclopedic materials) with dedicated long-context training up to 1M sequence length; post-training: supervised fine-tuning, large-scale reinforcement learning across capability domains, and MOPD consolidation of domain-specialized teacher policies. Trained on Huawei Ascend clusters. Supports 200+ languages.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
TAU3-Bench
general_agent
|
43.13%
|
— |
|
VITA-Bench
general_agent
|
53.89%
|
— |
|
SWE-bench Pro
coding_agent
|
55.50%
|
— |
|
SWE-bench Verified
coding_agent
|
47.25%
|
— |
|
SWE-bench Multilingual
coding_agent
|
57.16%
|
— |
|
AIME 26
stem_reasoning
|
89.16%
|
— |
|
HMMT Feb 26
stem_reasoning
|
78.63%
|
— |
|
IMOAnswerBench
stem_reasoning
|
73.30%
|
— |
|
Humanity's Last Exam
stem_reasoning
|
19.32%
|
— |
|
Workspace Bench
|
100.00%
|
— |
|
SciCode
|
84.91%
|
— |
|
Gaokao 2026
|
96.09%
|
— |
|
BFCL-V4
general_agent
|
83.10%
|
— |
|
TAU2-Bench
general_agent
|
72.84%
|
— |
|
BrowseComp
general_agent
|
42.91%
|
— |
|
IFEval
instruction_following
|
96.68%
|
— |
|
IFBench
instruction_following
|
87.02%
|
— |
|
AA-LCR
long_context
|
70.38%
|
— |
|
MCPMark
general_agent
|
12.91%
|
— |
|
GPQA
reasoning
|
84.24%
|
— |
|
MCP-Atlas
general_agent
|
57.46%
|
— |
Model artifacts and ecosystem
The repository contains model weights and configuration files for the post-trained model in Hugging Face Transformers format, distributed as Safetensors: 4B parameters with BF16 tensor type. The model uses custom code (custom_code tag), so pass --trust-remote-code when serving with vLLM. The base model is XHToken/Spark-X2.5-4B-Base; this checkpoint is a fine-tune of it. The ecosystem includes 5 public fine-tunes and 18 quantizations (usable with llama.cpp, LM Studio, Jan, and Ollama), and the models are also part of the Spark-X2.5 collection on Hugging Face.
Citation
If you find the work helpful, feel free to cite it:
@misc{sparkx2.5,
title = {Spark-X2.5 4B&1.7B: Pushing the Limits of Agentic Capabilities in On-Device Models},
author = {SparkLLM Team},
year = {2026}
}
License
The Spark-X2.5 model series is licensed under the Apache 2.0 License (see the LICENSE file in the model repository).
Recommended usage settings
- Sampling parameters: the recommended settings for Spark-X2.5 are
temperature=1.0,top_p=0.95, andtop_k=-1(example requests also userepetition_penalty=1,presence_penalty=0,frequency_penalty=0). - Thinking mode: thinking is enabled by default; to disable it for a single request, set
"chat_template_kwargs": {"enable_thinking": false}in the request body. - Long outputs: example requests use
max_tokensup to 131,072 for long-form generation. - Context length vs device memory: serving with the full 1,048,576-token context requires sufficient device memory; reduce
--context-lengthwhen necessary. - Fine-tuning: use LLaMA-Factory to customize the model.
- Agent harnesses: the models are deeply integrated with popular agent harnesses including Codex, Claude Code, OpenClaw, and Hermes.
Deployment: SGLang, vLLM, MLX, Ollama, LM Studio
Spark-X2.5 runs on a wide range of hardware — NVIDIA GPUs, Huawei Ascend NPUs (A2, A3, and Ascend 950DT), Apple silicon, and Linux CPU — and is compatible with leading inference frameworks such as vLLM, SGLang, llama.cpp, and MLX, plus quick deployment through Ollama and LM Studio.
SGLang (NVIDIA GPU). Use the pre-built image that tracks the Spark-X2.5 runtime (lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1; Ascend NPU daily builds are published under quay.io/ascend/sglang). Serve an OpenAI-compatible API with:
docker run --rm -it --gpus '"device=0"' --ipc=host -p 30000:30000 \
-v "$MODEL_PATH:/root/Spark-X2.5-4B:ro" \
lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1 \
python -m sglang.launch_server \
--model-path /root/Spark-X2.5-4B \
--served-model-name spark2.5 \
--tool-call-parser spark25 \
--reasoning-parser qwen3 \
--tp-size 1 --mem-fraction-static 0.8 \
--context-length 1048576 \
--chat-template /root/Spark-X2.5-4B/chat_template.jinja \
--host 0.0.0.0 --port 30000
Set MODEL_PATH to the absolute path of the local Spark-X2.5-4B checkpoint before starting the container.
vLLM (NVIDIA GPU). Use the official vllm/vllm-openai:latest image with --trust-remote-code and the model's chat_template.jinja (prefix caching optional):
docker run --rm --gpus all --ipc=host -p 30000:30000 \
-v "$MODEL_PATH:/models/Spark-X2.5-4B:ro" \
vllm/vllm-openai:latest \
--model /models/Spark-X2.5-4B --port 30000 \
--trust-remote-code --served-model-name spark25 \
--tensor-parallel-size 1 --gpu-memory-utilization 0.7 \
--enable-prefix-caching \
--chat-template /models/Spark-X2.5-4B/chat_template.jinja
For Ascend NPUs, use the official quay.io/ascend/vllm-ascend images (nightly-main for A2, nightly-main-a3 for A3, nightly-main-a5 for Ascend 950DT), then install the Spark plugin inside the container:
git clone https://github.com/XHToken/Spark-plugin.git
cd Spark-plugin
uv pip install .
MLX (Spark-MLX-LLM). Runs the original Spark-X2.5 Hugging Face checkpoints locally on Apple silicon GPU, Linux CPU, and NVIDIA CUDA on Linux — no GGUF conversion required:
git clone https://github.com/XHToken/Spark-MLX-LLM.git
cd Spark-MLX-LLM
python3 -m venv .venv && source .venv/bin/activate
python -m pip install -e . # Apple silicon (extras: '.[cpu]', '.[cuda12]', '.[cuda13]')
spark-mlx-generate --device gpu --dtype bfloat16 \
--model XHToken/Spark-X2.5-4B \
--prompt "What is the capital of Anhui Province?" \
--max-tokens 512 --temp 0
Ollama and LM Studio. Build the provided llama.cpp fork (https://github.com/XHToken/llama.cpp), then either create an Ollama model from a GGUF file (./ollama create Spark-X2.5-4B -f ./Modelfile.spark and ./ollama run Spark-X2.5-4B) or copy the llama.cpp-spark build output into an LM Studio runtime directory, place the GGUF model under the LM Studio models directory, and load it (or use the lms CLI: lms load <model>, lms chat <model>).
Fine-tuning. LLaMA-Factory is recommended for fine-tuning the model.
Thinking mode control
Thinking is enabled by default by both the chat template and the Qwen3 reasoning parser (the SGLang server is started with --reasoning-parser qwen3). To disable thinking for a specific request, set "chat_template_kwargs": {"enable_thinking": false} in the request body. All reported benchmark evaluations are conducted in thinking mode.
Benchmarks vs similarly sized models
Spark-X2.5 models are evaluated against leading on-device models of similar size across a broad range of tasks, including agent, code, math, and general & knowledge benchmarks.
| Benchmark | Spark‑X2.5‑4B | Spark‑X2.5‑1.7B | Qwen3.5‑9B | Qwen3.5‑4B | Qwen3.5‑2B | Gemma4‑12B | Gemma4‑E4B | Gemma4‑E2B |
|---|---|---|---|---|---|---|---|---|
| Agent | ||||||||
| BFCL‑V4 | 65.1 | 46.9 | 66.1* | 50.3* | 43.6* | 37.4 | 36.9 | 30.2 |
| τ²‑bench | 75.1 | 65.3 | 79.1* | 79.9* | 48.8* | 69.0* | 42.2* | 24.5* |
| τ³‑bench | 30.4 | 20.1 | 9.3 | 6.7 | 4.1 | 13.3 | 10.1 | 8.8 |
| MCP‑Atlas | 54.6 | 23.4 | 47.4* | 40.8* | 14.8 | 30.5* | 15.0* | 12.6 |
| MCP‑Mark | 14.2 | 2.3 | 13.4 | 12.5 | – | – | – | – |
| Workspace Bench | 31.2 | 18.9 | 25.5 | 21.3 | 7.7 | – | – | – |
| VitaBench2.0 | 25.2 | 8.3 | 15.6 | 18.2 | 5.2 | 12.4 | 4.8 | 4.4 |
| BrowseComp | 40.9 | 29.7 | 8.3 | 14.3 | 3.1 | 10.0 | 8.3 | 3.7 |
| Code | ||||||||
| SWE‑Bench Pro | 44.4 | 10.4 | 33.8* | 29.4* | 1.9 | 21.9* | 4.0* | – |
| SWE‑Bench Verified | 41.6 | 28.3 | 53.1* | 38.8* | 6.8 | 44.2* | 14.0* | – |
| SWE‑Bench Multilingual | 53.3 | 23.3 | 43.3 | 27.7 | 5.0 | 32.5* | – | – |
| SciCode | 34.7 | 18.2 | 32.7* | 24.0 | 6.0 | 39.8 | 27.5 | 20.5 |
| Math | ||||||||
| Gaokao 2026 | 133.4 | 114.8 | 135.5 | 130.3 | 94.0 | 130.6 | 102.4 | 81.8 |
| AIME 2026 | 90.7 | 69.4 | 88.2 | 83.0 | 30.8 | 82.1* | 42.5* | 37.5* |
| HMMT Feb 2026 | 81.2 | 48.4 | 70.8 | 69.7 | 21.5 | 65.6 | 34.2 | 20.5 |
| IMO‑AnswerBench | 74.2 | 45.4 | 69.8 | 68.5 | – | 57.2 | 26.9 | 22.6 |
| General & Knowledge | ||||||||
| IFEval | 93.0 | 89.5 | 91.5* | 89.8* | 78.6* | 94.8 | 45.3 | 34.8 |
| IFBench | 75.0 | 66.3 | 64.5 | 59.2 | 41.3* | 73.5* | 44.0* | 22.7 |
| AA‑LCR | 56.3 | 24.3 | 63.0* | 57.0* | 25.6* | 55.3* | 34.7 | 18.3 |
| HLE | 12.3 | 6.3 | 14.3 | 8.6 | 2.1 | 13.1 | 3.9 | 2.5 |
| GPQA | 67.4 | 43.8 | 77.2 | 67.2 | 44.6 | 72.8 | 54.5 | 43.8 |
Notes:
*denotes reported results from publicly released model cards / papers;-denotes scores not yet available.- All evaluations are conducted in thinking mode. The recommended sampling parameters for Spark-X2.5 are temperature=1.0, top_p=0.95, and top_k=-1.
- Gaokao 2026 consists of the five 2026 Chinese GAOKAO examinations (National I, National II, Beijing, Shanghai, Tianjin), each graded out of 150 points.
Training methods: pretraining data and post-training
Pretraining. Spark-X2.5 is pretrained on approximately 20 trillion tokens from a diverse corpus spanning web pages, books, academic publications, code, and encyclopedic materials. Particular attention is paid to data quality, domain coverage, and the sampling weights assigned to different data categories. Extensive data-mixture studies determine an effective balance among mathematics, logic, code, and other high-value domains, enabling the models to acquire broad general knowledge while developing stronger capabilities in complex reasoning and code generation. Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens.
Post-training. Post-training begins with supervised fine-tuning (SFT) on a carefully curated corpus, establishing robust instruction following, structured generation, and task completion while providing a stable policy initialization for reinforcement learning. Large-scale reinforcement learning is then applied across several capability domains — language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following — yielding a set of domain-specialized teacher policies whose complementary strengths are consolidated into a single deployable model through MOPD. The models were trained on Huawei Ascend clusters.
Native 1M-token context window
Spark-X2.5 natively supports a context window of up to 1M tokens (1,048,576). Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens. When serving with SGLang, start the server with --context-length 1048576; this setting requires sufficient device memory, so reduce --context-length when necessary. Example client requests generate up to 131,072 tokens (max_tokens: 131072).
Hybrid attention architecture
For agent tasks, balancing performance, inference speed, and cache usage has long been a key bottleneck limiting model performance. Spark-X2.5 systematically integrates and optimizes mature attention technologies, combining sliding-window attention (SWA) with a hybrid full-attention architecture: one full-attention layer is paired with three sliding-window attention layers. This design leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure, achieving an effective balance among performance, inference efficiency, and KV-cache size — thereby improving practicality and effectiveness in real-world deployment scenarios.
Introduction and technical highlights
Spark-X2.5-4B and Spark-X2.5-1.7B are compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. They deliver strong performance across everyday tasks — conversation, writing, translation, reasoning, coding, tool use, and agentic workflows — achieving leading results among open-source models of comparable size. They combine an efficiency-oriented architecture with native context windows of up to 1M tokens and support for more than 200 languages.
Technical highlights
- Efficient architecture and native 1M-token context: a hybrid attention architecture combines one full-attention layer with three sliding-window attention layers, substantially reducing the computational overhead of long-context models while natively supporting a context window of up to 1M tokens.
- Strong coding and agent capabilities: deeply integrated with popular agent harnesses including Codex, Claude Code, OpenClaw, and Hermes; state-of-the-art performance among models of comparable size across everyday coding, agentic workflows, reasoning, and instruction following.
- Broad hardware and software compatibility: supports NVIDIA, Huawei, Hygon, HOUMO.AI, and other hardware platforms; compatible with leading inference frameworks such as vLLM, SGLang, llama.cpp, and MLX; deployable quickly through Ollama and LM Studio; customizable with fine-tuning frameworks such as LLaMA-Factory. Delivers superior TTFT, TOPT, and overall inference efficiency compared with similarly sized models.
- Advanced training algorithms: trained on Huawei Ascend clusters; large-scale reinforcement learning and post-training techniques such as MOPD significantly enhance reasoning, coding, agentic, and instruction-following capabilities.
Architecture
- Attention
- Hybrid Attention (16:4)
- Layers
- 36
- Hidden size
- 2560
- Context
- 1M tokens
- Parameters
- 4000M
Source: Hugging Face config.json · Spark2_5ForCausalLM · exact layer pattern · model repo
Training Pipeline
-
1
pretraining
Pretraining (~20T tokens)
Pretrained on approximately 20 trillion tokens from web pages, books, academic publications, code, and encyclopedic materials, with data-mixture studies balancing mathematics, logic, and code; long-context capability developed in a dedicated stage with sequence lengths up to 1M tokens.
-
2
sft
Supervised fine-tuning
Supervised fine-tuning on a carefully curated corpus establishing instruction following, structured generation, and task completion, providing a stable policy initialization for reinforcement learning.
-
3
rl
Large-scale reinforcement learning
Large-scale RL across capability domains including language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following, yielding domain-specialized teacher policies.
-
4
other
MOPD policy consolidation
MOPD consolidates the complementary strengths of the domain-specialized teacher policies into a single deployable model.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| ~20T token pre-training corpus (web, books, academic, code, encyclopedic) | pretraining | — | — | |
| SFT + large-scale RL + MOPD consolidation corpora | rl | — | — |
Linked Resources
SparkLLM Slack
https://join.slack.com/t/tokenspark/shared_invite/zt-432qf8l2f-5~dLyXv8uETr0P0UuC07nw
SparkLLM Discord
https://discord.gg/kTDE2Hg8aw
SparkLLM YouTube
https://www.youtube.com/@SparkLLM
SparkLLM on dev.to
https://dev.to/sparkllm
SparkLLM on Bluesky
https://bsky.app/profile/sparkllm.bsky.social
SparkLLM on X (Twitter)
https://x.com/sparkllm
SparkLLM on Zhihu
https://www.zhihu.com/people/zhiikz7qh7m
Spark-plugin (vLLM/Ascend plugin)
https://github.com/XHToken/Spark-plugin.git
Spark-MLX-LLM (MLX runtime)
https://github.com/XHToken/Spark-MLX-LLM.git
llama.cpp (Spark fork for GGUF/Ollama/LM Studio)
https://github.com/XHToken/llama.cpp.git
LlamaFactory (Spark fine-tuning fork)
https://github.com/XHToken/LlamaFactory
Spark-X2.5 Hugging Face Collection
https://huggingface.co/collections/XHToken/spark-x25